Open samples
Judge the data before you talk to us.
Buyers judge a data vendor by its samples, and most keep theirs behind a sales call. Ours are here: browse every record, see how it was checked, and download it. They were written by our team to show the format and the bar; client data never appears here.
RL environment
A practice copy of a company's expense-review app with a written policy. Agents work the queue; every task is scored from the final state.
50 scored tasks
- 13 tools
- Currencies, duplicates, split purchases, budgets
- Injected instructions in employee notes
- Playable in the browser
✓50/50 verified: reference passes, 4 grader attacks fail
Verified coding tasks
Bug fixes and features as an agent would receive them: problem, starting code, hidden tests and a reference solution.
12 tasks
- Python, TypeScript, JavaScript
- 33 fail→pass tests
- SQL injection and prototype pollution fixes
✓12/12 verified in the sandbox: tests fail before the fix, pass twice after
Preference data
Two AI responses, the better one, how strongly, and why. Includes agent safety, honesty and a genuine tie.
12 pairs
- Coding and agent behavior
- Choice + strength + rationale
- Raw and DPO formats
✓Every pair has a choice, a strength and a rationale that names the deciding issue (checked by tests)
Rubric grading
Finance and accounting answers scored per criterion, with a written reason for every grade.
8 graded answers
- Accuracy, completeness, clarity
- US GAAP and ASC 606
- Wrong core claims capped at 2
✓Every grade follows the rubric's cap rule (checked by tests)
Want this for your domain?
Most pilots start with 50–200 tasks, and you pay only for approved answers.