Canonset

Open samples

Judge the data before you talk to us.

Buyers judge a data vendor by its samples, and most keep theirs behind a sales call. Ours are here: browse every record, see how it was checked, and download it. They were written by our team to show the format and the bar; client data never appears here.

RL environment

Expense approvals environment

A practice copy of a company's expense-review app with a written policy. Agents work the queue; every task is scored from the final state.

50 scored tasks

  • 13 tools
  • Currencies, duplicates, split purchases, budgets
  • Injected instructions in employee notes
  • Playable in the browser

50/50 verified: reference passes, 4 grader attacks fail

Verified coding tasks

Verified coding tasks

Bug fixes and features as an agent would receive them: problem, starting code, hidden tests and a reference solution.

12 tasks

  • Python, TypeScript, JavaScript
  • 33 fail→pass tests
  • SQL injection and prototype pollution fixes

12/12 verified in the sandbox: tests fail before the fix, pass twice after

Preference data

Preference comparisons

Two AI responses, the better one, how strongly, and why. Includes agent safety, honesty and a genuine tie.

12 pairs

  • Coding and agent behavior
  • Choice + strength + rationale
  • Raw and DPO formats

Every pair has a choice, a strength and a rationale that names the deciding issue (checked by tests)

Rubric grading

Rubric grading

Finance and accounting answers scored per criterion, with a written reason for every grade.

8 graded answers

  • Accuracy, completeness, clarity
  • US GAAP and ASC 606
  • Wrong core claims capped at 2

Every grade follows the rubric's cap rule (checked by tests)

Want this for your domain?

Most pilots start with 50–200 tasks, and you pay only for approved answers.