Senior Application Security Engineer - AI Code Evaluation
Remote
Leading AI Lab · Paris, FR / US
Fresh
Gig
Braintrust
Compensation
$75-$90/hr
Description
Help a top AI lab evaluate and improve large language models through security-focused coding tasks. Bring your software-engineering judgment and hands-on security experience to work involving vulnerabilities, exploit verification and security patches.
This is a contracting engagement, with potential for a longer-term engagement. Remote candidates in the selected countries are elegible.
What you will do
- Evaluate coding tasks involving software vulnerabilities, exploit verification and security patches.
- Create high-quality coding prompts and reference answers for benchmark-style problems.
- Evaluate model outputs for code generation, refactoring, debugging and implementation.
- Identify and document model failures, edge cases and reasoning gaps.
- Compare private language models with leading external models.
- Build or configure coding environments for evaluation and reinforcement learning.
- Follow detailed annotation and evaluation guidelines consistently.
What you bring
- At least five years of professional software-development experience and strong Python skills.
- Hands-on experience with vulnerability research, exploit reproduction or verification, or implementing, backporting or validating security patches.
- The ability to apply structured evaluation criteria and write clear technical feedback.
- Fluency in written and spoken English.
Helpful, not required
- Professional code review, coding annotation, LLM/code evaluation or benchmark design.
- Knowledge of another programming language.
- Team leadership or mentoring experience.
- Contract
- Long
- Engagement
- Freelance
Skills & categories
LLMsReinforcement LearningPythonSoftware EngineeringData Annotation
- Posted
- Sep 28, 2026
- Slots remaining
- 10
- First seen
- Sep 28, 2026
- Last seen
- Sep 29, 2026