Kairos
Back to gigs

Security Software Engineer - Python & AI Evaluation

Remote

Leading AI Lab · Paris, FR / US

Fresh
Gig
Braintrust

Compensation

$75-$90/hr

Apply on Braintrust

Description

Help a top AI lab evaluate and improve large language models through security-focused coding tasks. Bring your software-engineering judgment and hands-on security experience to work involving vulnerabilities, exploit verification and security patches.

This is a contracting engagement, with potential for a longer-term engagement. Remote candidates in the selected countries are elegible.

What you will do

  • Evaluate coding tasks involving software vulnerabilities, exploit verification and security patches.
  • Create high-quality coding prompts and reference answers for benchmark-style problems.
  • Evaluate model outputs for code generation, refactoring, debugging and implementation.
  • Identify and document model failures, edge cases and reasoning gaps.
  • Compare private language models with leading external models.
  • Build or configure coding environments for evaluation and reinforcement learning.
  • Follow detailed annotation and evaluation guidelines consistently.

What you bring

  • At least five years of professional software-development experience and strong Python skills.
  • Hands-on experience with vulnerability research, exploit reproduction or verification, or implementing, backporting or validating security patches.
  • The ability to apply structured evaluation criteria and write clear technical feedback.
  • Fluency in written and spoken English.

Helpful, not required

  • Professional code review, coding annotation, LLM/code evaluation or benchmark design.
  • Knowledge of another programming language.
  • Team leadership or mentoring experience.


Contract
Long
Engagement
Freelance

Skills & categories

LLMsReinforcement LearningPythonSoftware EngineeringData Annotation
Posted
Sep 28, 2026
Slots remaining
10
First seen
Sep 28, 2026
Last seen
Sep 29, 2026