Research Intern

About Scorecard

We're a small team within Scorecard, the simulation platform for AI agents, and we work closely with our customers to build multi-user simulations.

The role

This team focuses on creating simulations for agents interacting with multiple users, which means we need simulated users (“simusers”) to interact with the agent. Currently, engineers simulate users using vibecoded prompts like “act frustrated,” resulting in unrealistic users that corrupt evaluation results.

To fix this, we’re building a simuser framework and orchestration layer. It allows defining high-fidelity simusers and realistically managing a group of simusers that interact with the agent-under-test in a simulation.

We're hiring a research intern to own one question in this space end-to-end and publish what they find, paired with a dedicated mentor from the team throughout. Internships run 10 to 16 weeks with flexible start dates throughout the year.

What you'll do

  • Own a research question end to end. Literature, experiment design, prototype, evaluation, write-up. We scope the project to your strengths when you arrive. Example questions: how do we measure that a simuser behaves like the specific person it was built from? What can we recover about someone's goals, knowledge, and style from their work transcripts? How can we avoid hallucination and sim drift in long-horizon simulations?
  • Work with real data. Our grounding data is human work transcripts. Expect to spend time cleaning it and to develop opinions about what it can and can't tell you.
  • Ship into the framework. Your prototypes land in the codebase, and your experiments run against the same system Scorecard's platform and partners use.
  • Publish. You'll be an author on the work you ship. We expect strong results to become a paper or public benchmark, and we write up intern projects on our blog.

What we're looking for

  • Research experience. You're an MS or PhD student, or have done comparable research outside a degree program. You've turned vague questions into rigorous experiments and know what makes a result trustworthy.
  • LLM experience. You've built with language models before, e.g. evals, agent scaffolding, or data pipelines, and can run experiments with Python.
  • A feel for agent behavior. You notice when a simulated user sounds off, and you can explain why and check whether a fix worked.
  • Hands-on ownership. The framework is new and the team is small. You're comfortable making design decisions and experimenting with incomplete information.

These aren’t hard requirements and you don't need publications. If you've done solid research and this problem interests you, apply!

Nice to have

  • Published research: papers, benchmarks, datasets, or public experiments.
  • Prior work on user simulation: simulated users for dialogue or agent benchmarks, digital twins, role-play systems, or character consistency.
  • Evaluation work: LLM-as-judge design, grader calibration, or human annotation pipelines.
  • RL experience: building RL environments for LLMs, reward design, or firsthand encounters with reward hacking.
  • Behavioral or social-science research methods: survey design, test-retest protocols, inter-rater reliability.

Compensation and benefits

  • In-office perks and daily lunch
  • Competitive hourly rate + housing stipend

How to apply

Interested? Send your resume and a quick note to jobs@scorecard.io.

Location
Employment Type
Team

Take your first step towards self-improving agents

Join forward-thinking teams using Scorecard to upgrade the way they build, test, and improve AI AGENTS.

Learn More