Compensation
Salary undisclosedDescription
As a Staff AI Engineer at Factored, you will operate at the intersection of technical architecture, Generative AI, and enterprise strategy. Working directly with enterprise clients, you will act as a key technical contributor and trusted advisor.
This role is for technical leaders who combine system architecture mastery and hands-on ML/GenAI engineering with the executive presence needed to navigate ambiguous environments. You will define business problems, architect production-grade AI applications, align senior stakeholders, and own end-to-end delivery to drive measurable impact.
Functional Responsibilities:
- Partner with client executives to translate ambiguous business problems into enterprise AI solution architectures with clear trade-off analyses (cost, latency, risk).
- Design and build scalable backend systems, data pipelines, and APIs integrating LLMs, agentic workflows,and RAGs.
- Implement multi-agent orchestration frameworks and advanced retrieval mechanisms (vector DBs, hybrid search) for complex workflows.
- Deploy and manage cloud-native AI applications across AWS, GCP, Azure, or Databricks using Docker, Kubernetes, Terraform, and CI/CD pipelines.
- Instrument systems with LLM telemetry, cost-tracking, security guardrails, and systematic evaluation harnesses (LLM-as-a-judge) to ensure safety and performance.
- Fine-tune prompts and optimize inference latency using caching, quantization, and cost-reduction strategies.
- Serve as the embedded technical authority within client environments to align cross-functional teams and manage technical risks.
- Elevate team standards (modular code, testing, CI/CD) and mentor client technical staff to build long-term operational autonomy
Qualifications:
- 8+ years of experience in Software/ML Engineering, with 3+ years specifically focused on production GenAI/LLM applications (RAG, agents, tool use) and 2+ years in customer-facing or forward-deployed roles.
- Deep hands-on experience building production systems with Generative AI frameworks (LangGraph, LangChain, LlamaIndex, OpenAI, vector databases).
- Proven ability to architect and scale complex backend microservices and APIs using Python (FastAPI, Django, Flask) alongside relational and NoSQL databases.
- Hands-on expertise building, deploying, and managing cloud-native applications on AWS, GCP, Azure, or Databricks using Docker, Kubernetes, Terraform, MLflow, and automated CI/CD pipelines.
- Experience implementing LLM telemetry, cost-tracking, security guardrails, and systematic evaluation harnesses (LLM-as-a-judge patterns).
- Exceptional ability to structure ambiguous client problems into clear technical requirements and present trade-off analyses (cost, latency, risk) to non-technical executive stakeholders.
- Fluent English communication (written and spoken) with a track record of driving engagements independently in fast-paced, high-stakes environments.
Stack
LLMsGenerative AIPythonLangGraphGCPAzureTerraformCI/CDLangChainLlamaIndexAgentic AIVector DatabasesAWSMachine LearningKubernetesDockerDatabricksFastAPIDjangoRAGFlaskData EngineeringMLflowQuantization
- Posted
- Sep 16, 2026
- Last seen
- Sep 17, 2026
- First seen
- Sep 17, 2026


