Kairos
Back to jobs

Sr. Staff Inference Serving Engineer

On-site
GroqDallas, TX, US3 hours agoWebsite
Fresh
Full-time
Staff / Principal
Platform Eng

Compensation

Salary undisclosed
Apply
Share

Description

About Groq
Inference is the engine that powers AI, and Groq was built from the silicon up to deliver the world's fastest inference at scale. We pioneered the LPU—the first processor designed specifically for AI inference—and are transforming that innovation into a global cloud platform powering production AI workloads.

With the capital, infrastructure, and team to execute, we're uniquely positioned to define the next era of AI infrastructure. The opportunity is massive, and it's still wide open. Now let's go build it!
Mission:
GroqCore is building a high-performance software stack that converts bare metal compute into a token-producing engine This role is for a seasoned engineer with practical experience making LLMs fast, efficient, and reliable in production. This individual will work on an inference serving software stack from a trained model to a high-throughput delivery of output tokens. This will require experience on understanding/working with model internals, runtime software, and the supporting GPU & accelerator hardware underpinning it all.

Location: This role will be based in one of our three hiring hubs: the Dallas, San Francisco, or New York City area. The person hired for this role must be based in one of these three areas. You’ll have the flexibility to work remotely while we establish our local Groq office, with the expectation that this role will transition to onsite once the office opens.

Responsibilities & opportunities in this role:
  • Build and optimize serving systems that handle high request volumes with low, predictable latency
  • Apply model compilation and graph optimization techniques, and benchmark carefully to use them only where they deliver real gains
  • Implement and tune batching, caching, scheduling, and memory management strategies for large models
  • Apply reduced-precision and quantization methods 
  • Distribute models across multiple accelerators and nodes, balancing throughput & latency
  • Profile end-to-end performance, identify bottlenecks, and fix them—ranging from low-level kernels to request routing—leveraging both lab/development environments and live telemetry from production systems.
  • Partner with infrastructure and FDE teams to bring new models into production quickly
  • Partner with vendors along the inference serving path in support of optimizing the stack
Ideal candidates have/are:
  • BS / MS / PhD in CS, CE, EE, or equivalent depth from industry
  • 5+ years shipping performance-critical products, with some portion of that serving models at scale
  • A strong applied understanding of transformer architectures
  • Fluency in at least one systems language and one high-level language
  • A rigorous, measurement-driven approach to performance work
  • Clear communication about tradeoffs to both technical and business stakeholders
Ways to stand out:
  • Experience writing or tuning custom kernels.
  • Familiarity with non-GPU or specialized inference hardware.
  • Contributions to open-source serving or compiler projects.
Why Join Us:
  • Purposeful Hiring: You’re not here by accident, and neither is anyone else. Every teammate is handpicked with intention because who we build with matters.
  • Builders Wanted: You’re not just riding the rocket ship, you’re building it. Your work directly shapes the trajectory of our company.
  • Mission-Driven Work: We’re here to make a real impact. Our mission fuels everything we do.
  • Tackling Hard Problems: If easy isn’t your thing, you’re in the right place. We solve some of the most complex and exciting challenges in our space.
  • Excellence Is The Standard: High performance isn’t just encouraged, it’s the baseline. And it’s contagious.

If this sounds like you, we’d love to hear from you!

Compensation
Groq is committed to providing competitive compensation through our Total Cash philosophy, which incorporates potential bonus value directly into base pay. The total cash salary range for this position, which is inclusive of the potential bonus value, is $341,400 - $401,600, with individual placement determined by your geographic location, experience, skills, and alignment with internal compensation standards. This range is specific to candidates located in the United States. Compensation for international candidates will vary based on local market dynamics. Beyond cash compensation, Groq also offers a Long-Term Incentive (LTI) Program and a robust suite of employee benefits.
#LI-MS1
US Job Posting
This position may require access to technology and/or information subject to U.S. export control laws and regulations, including the Export Administration Regulations (EAR). To comply with these requirements, candidates for this role must meet certain citizenship or residency criteria. Specifically, they must qualify as U.S. Persons for export control purposes (i.e., U.S. citizen, U.S. lawful permanent resident (Green Card holder), or a protected individual under 8 U.S.C. § 1324b(a)(3) such as a refugee or asylee), or otherwise be eligible for an applicable export license.

Non-US Job Postings
This position may require access to technology and/or information subject to U.S. export control laws and regulations, as well as applicable local laws and regulations, including the Export Administration Regulations (EAR). To comply with these requirements, candidates for this role must meet all relevant export control eligibility criteria.

Groq is an Equal Opportunity Employer. We are committed to creating an inclusive environment for all employees and applicants. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, sex (including gender identity, sexual orientation, and pregnancy), age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable law.

Groq complies with all applicable federal, state, and local laws governing nondiscrimination in employment. We do not tolerate discrimination or harassment based on any protected characteristic.

Groq is committed to working with and providing reasonable accommodations to qualified individuals with physical or mental disabilities. If you require a reasonable accommodation to complete an application or to participate in the hiring process, please contact us at talent@groq.com. This contact is for accommodation requests only, which will be considered on a case-by-case basis.

All offers of employment are contingent upon verification of the applicant’s identity and employment authorization in accordance with federal law.

Groq encourages people with criminal record histories to apply for employment, and values diverse experiences, including prior contact with the criminal legal system.  To that end, Groq welcomes such applicants in accordance with the California Fair Chance Act, Los Angeles City Fair Chance Act Ordinance, Los Angeles County Fair Chance Act Ordinance, and San Francisco Fair Chance Act Ordinance.  Philadelphia applicants can review information pertaining to Philadelphia’s Fair Criminal Record Screening Standards Ordinance here: https://www.phila.gov/documents/fair-chance-hiring-law-poster.

As part of our hiring process, Groq may use artificial intelligence (“AI”) tools or automated systems to assist with activities such as reviewing applications, evaluating qualifications, scheduling interviews, analyzing assessment responses, or supporting recruiting operations. These tools are designed to assist—not replace—human decision-making, and hiring decisions are subject to human review. We may process information you provide during the application process, including resumes, application materials, interview responses, assessments, and, where applicable, audio, video, or transcript data. If legally required, we will request consent before using technologies that analyze biometric or video interview data. Candidates may request reasonable accommodations, an alternative evaluation process, additional information regarding the use of AI in the hiring process, or review of certain automated decisions by contacting ta-operations@groq.com.

Stack

GPULLMsQuantization
Posted
Oct 9, 2026
Last seen
Oct 9, 2026
First seen
Oct 9, 2026

Similar roles

Browse more AI jobs