Skip to sign up

Job matched to your search

C

Research Engineer, Benchmarks

Clera · Singapore

Singapore · RemoteFull-TimeUSD 150k - USD 250kPosted Sep 4, 2026

Free · Join 5,000+ job seekers using Qarera

How well do you match this role?

Tap the skills you already have — then see your real match score, what’s missing, and your resume fixed for this job.

↑ tap the skills you have
Loading sign-in…
Free · no credit card · 30 seconds

Job description

ABOUT THE ROLE

You will own the design and implementation of rigorous, domain-specific benchmarks used to evaluate frontier AI agents on realistic workflows. Sitting within a small, highly technical team of researchers and engineers, this role is central to delivering evaluations that AI labs and enterprise customers genuinely trust and rely on.

WHAT YOU'LL DO

  • Design, implement, and maintain high-quality internal benchmarks for evaluating frontier agents on domain-specific tasks.

  • Partner with subject-matter experts to define realistic workflows and translate them into well-scoped evaluation tasks.

  • Build and operate reliable infrastructure to run models and agents against benchmark tasks at scale.

  • Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.

  • Validate that benchmark performance correlates with real-world evaluations and customer needs.

  • Write clear technical documentation and benchmark reports for research and engineering audiences.

WHAT WE'RE LOOKING FOR

  • 2 to 4 years of experience in research engineering or ML engineering, with a focus on AI benchmarks, evaluation infrastructure, or agent environments.

  • Strong proficiency in Python, Docker, and Linux for building research or production infrastructure.

  • Demonstrated experience designing and running benchmarks or evaluation environments for AI agents or large language models.

  • Experience building infrastructure to reliably run AI models or agents against evaluation tasks at scale.

  • Experience developing metrics or validation studies to assess benchmark difficulty, reliability, and real-world correlation.

  • Ability to collaborate with domain experts and translate complex workflows into evaluation criteria.

  • Strong attention to detail, with a habit of spotting subtle inconsistencies and edge cases.

  • Comfort working independently in fast-paced, early-stage startup environments with unstructured problem spaces.

  • Excellent written communication skills for technical documentation and cross-timezone collaboration.

  • Published papers or technical writing on AI benchmarking, model evaluation, or failure modes is a strong plus.

  • Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation is a plus.

COMPENSATION & BENEFITS

Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.

LOCATION

On-site in Singapore.

Don’t just read the job — see if you’ll get it.

Get your match score, a resume tailored to this exact role, and jobs like it — free.

Check my fit for this job
Loading sign-in…
Apply →

Hiring for a role like this? Join the employer waitlist.