Cloud Geometry
Senior AI/ML Engineer
Descripción
About the Role
We're looking for a Senior AI/ML Engineer to join our flagship AI Platform for the life sciences industry, supporting global leaders like Pfizer, Moderna, and Novartis in accelerating drug discovery and RNA-based innovation through cloud and AI.
You'll architect and deploy production-grade AI systems, from autonomous agent frameworks and MCP-powered tool orchestration to classical ML pipelines and enterprise AI gateways. This is a senior role with real ownership: challenge architectural decisions, propose improvements, and own features end to end.
What You'll Do
Build and deploy multi-agent AI systems (LangGraph, AutoGen, CrewAI, or custom loops) using tool-calling and ReAct-style reasoning, and implement MCP (Model Context Protocol) servers and clients that expose tools, resources, and prompts to LLM agents. Design agent memory — short-term (in-context) and long-term (vector store / KV) — and human-in-the-loop approval gates for regulated life-sciences workflows. Deploy AI Gateway infrastructure to manage LLM routing, fallback, cost, and security across providers (OpenAI, Anthropic, Bedrock, Vertex), including PII redaction, prompt-injection guards, audit logging, and RBAC for GxP-adjacent environments. Build end-to-end ML training and deployment pipelines (XGBoost, PyTorch, Databricks/SageMaker) with experiment tracking (MLflow), containerized serving on ECS/EKS, and monitoring for drift and performance. Design and optimize RAG pipelines — chunking, embeddings, vector stores (pgvector, OpenSearch, Pinecone), re-ranking — and fine-tune open-source LLMs (LoRA/QLoRA) for domain-specific use cases. Build structured prompt libraries and LLM-as-judge evaluation with red-team / adversarial testing. Lead architecture reviews, mentor on MLOps best practices, and collaborate with data engineers on Databricks lakehouse pipelines. Participate in daily Scrum with a globally distributed team.
What You Bring
5+ years in software/AI engineering, including 3+ years focused on AI/ML in production. 2+ years deploying classical ML models and/or LLM-based systems at scale. Hands-on experience with AI agent frameworks and/or AI Gateway infrastructure (Portkey, Bedrock Gateway, Kong AI Gateway, LiteLLM, or custom). Direct experience implementing MCP servers/clients and tool/function-calling APIs with structured outputs (JSON Schema, Zod, Pydantic). Strong Python skills; working knowledge of TypeScript/Node.js for backend APIs (FastAPI, Express/HapiJS). Experience with Databricks (Spark, Delta Lake, MLflow, Unity Catalog) and AWS (ECS, Lambda, SageMaker, S3, plus SQS/SNS, Step Functions, OpenSearch). Familiarity with agent/ML observability (LangSmith, Arize Phoenix, or OpenTelemetry) and model monitoring (Evidently, WhyLabs). Excellent English communication — able to explain ML systems to both technical and non-technical stakeholders. Comfortable working autonomously in a remote team (9 AM – 5 PM EST overlap required). AWS ML Specialty, Databricks ML Professional, or Google ML Engineer certification. Experience fine-tuning open-source LLMs (Llama, Mistral, Falcon). Familiarity with GxP / 21 CFR Part 11 compliance in life sciences. Contributions to open-source AI/ML or MCP tooling. Experience building internal ML platforms, developer SDKs, or self-service ML tooling.
What We Offer
Competitive compensation and benefits. Zero legacy infrastructure — work on cutting-edge AI systems. Training budget (certifications, hackathons, Udemy/Coursera). Access to Claude Code, Codex, Cursor, and frontier model APIs. A collaborative team at the intersection of life sciences and applied AI.
We want to be upfront: we may use AI tools to help our team review applications and assess responses. These tools are there to support our recruiters, not replace them. Every hiring decision comes down to a human. If you have questions about how your information is used, we're happy to help.