Research Engineer
About Scorecard
We’re a small, nimble team backed by top-tier investors. We have multi-billion dollar customers and a long runway. To accelerate our growth, we’re looking to hire exceptional engineers to help shape the future of AI development.
This role is part of the "simuser research" pod within Scorecard, and we work closely with our customers to build multi-user simulations.
The role
This team focuses on creating simulations for agents interacting with multiple users, which means we need simulated users (“simusers”) to interact with the agent. Currently, engineers simulate users using vibecoded prompts like “act frustrated,” resulting in unrealistic users that corrupt evaluation results.
To fix this, we’re building a simuser framework and orchestration layer. It allows defining high-fidelity simusers and realistically managing a group of simusers that interact with the agent-under-test in a simulation.
We’re hiring a research engineer to help us build and evaluate this framework.
What you'll do
- Build the simuser orchestrator. A simulation containing multiple simusers has to manage conversations and context between simusers. An orchestrator is necessary to prevent simulation drift and keep simusers realistic.
- Build the simuser framework. We need a structured schema to define the simuser’s goals, knowledge, and style, and a tool harness for the simusers to take action in the environment.
- Convert transcripts into simusers. Our simusers start from transcripts of a real human doing the task. You'll build the pipeline that turns work context and transcripts into a digital twin that replicates that person's goals, knowledge, quirks, and memory.
- Design and run evals. We need to validate that our approach to creating simusers results in sufficiently realistic simulations. We plan to publish a benchmark to quantify the fidelity of a grounded simulated user.
- Ship to production. Scorecard's platform consumes everything you build; your simusers run inside training environments where their failures become the agent’s habits.
What we're looking for
- Engineering fundamentals. You write production Python or TypeScript and have owned a system from design through deploy.
- LLM fluency. You've built agentic loops or multi-agent systems and iterated on them with an eval in the loop.
- Empirical rigor. Fidelity claims need evidence. You've designed evaluations with baselines, held-out data, and calibrated judges, and you notice when a metric is easy to hack.
- Comfort with messy data. Our simusers are grounded in real human data. You've worked with data like this and know how much cleaning and judgment it takes before anything structured comes out of it.
- A feel for agent behavior. You can read a simulated conversation and tell when the user is off, like being too articulate or too helpful. You can turn that observation into a rubric or a prompt change, and evaluate it on whether it actually helped.
- Hands-on ownership. The framework is new and the team is small. You're comfortable making design decisions with incomplete information, building the first version yourself, and revising it when the data disagrees with you.
Nice to have
- Published research: papers, benchmarks, datasets, or public experiments.
- Prior work on user simulation: simulated users for dialogue or agent benchmarks, digital twins, role-play systems, or character consistency.
- Evaluation work: LLM-as-judge design, grader calibration, or human annotation pipelines.
- RL experience: building RL environments for LLMs, reward design, or firsthand encounters with reward hacking.
- Behavioral or social-science research methods: survey design, test-retest protocols, inter-rater reliability.
Compensation and benefits
- In-office perks and daily lunch
- Medical, dental, and vision benefits
- Unlimited PTO
- Annual salary range: $175,000 - $250,000 + competitive equity stake
How to apply
Interested? Send your resume and a quick note to jobs@scorecard.io.