Canonset
← Samples

Verified coding tasks

Tasks that prove their own tests

Each task is what a coding agent would receive: a problem statement and starting code, plus hidden tests and a reference solution. Our sandbox ran every one: the tests fail on the starting code, pass twice on the reference, and have no network access.

  • 12 tasks
  • Python 3.12, TypeScript, JavaScript
  • 33 fail→pass tests
  • Bug fixes, features, two security fixes

Split a bill in cents without losing a cent

code-12 · what the agent sees is the problem statement and the starting code

easyJavaScript · Node.js 24 · node:test
Problem statement

split.js divides a bill between people by weight (for example nights stayed). Finance found that splits often add up to a cent more or less than the bill. Rewrite splitCents(totalCents, weights) so that it returns whole cents that always add up to totalCents exactly. Give each person the floor of their exact share, then hand out the remaining cents one at a time to the largest fractional remainders, earliest person first on ties. Throw RangeError when totalCents is not a non-negative integer, weights is empty, or any weight is negative, or all are zero.

Starting code
split.js
export function splitCents(totalCents, weights) {
  const sum = weights.reduce((a, b) => a + b, 0);
  return weights.map((w) => Math.round((totalCents * w) / sum));
}
Reference solution
split.js
export function splitCents(totalCents, weights) {
  if (!Number.isInteger(totalCents) || totalCents < 0) throw new RangeError("totalCents must be a non-negative integer");
  const sum = weights.reduce((a, b) => a + b, 0);
  if (!weights.length || sum <= 0 || weights.some((w) => w < 0)) throw new RangeError("weights must be non-negative and not all zero");
  const exact = weights.map((w) => (totalCents * w) / sum);
  const shares = exact.map(Math.floor);
  let left = totalCents - shares.reduce((a, b) => a + b, 0);
  const order = exact.map((x, i) => ({ i, remainder: x - Math.floor(x) })).sort((a, b) => b.remainder - a.remainder || a.i - b.i);
  for (const { i } of order) {
    if (left === 0) break;
    shares[i] += 1;
    left -= 1;
  }
  return shares;
}
Tests (hidden from the agent)
split.test.js
import assert from "node:assert/strict";
import { test } from "node:test";
import { splitCents } from "./split.js";

test("even split", () => {
  assert.deepEqual(splitCents(900, [1, 1, 1]), [300, 300, 300]);
});

test("weighted split", () => {
  assert.deepEqual(splitCents(1001, [2, 1]), [667, 334]);
});

test("always adds up to the bill", () => {
  assert.deepEqual(splitCents(1000, [1, 1, 1]), [334, 333, 333]);
});

test("remainders go to the earliest people on ties", () => {
  assert.deepEqual(splitCents(100, [1, 1, 1, 1, 1, 1]), [17, 17, 17, 17, 16, 16]);
});

test("rejects bad input", () => {
  assert.throws(() => splitCents(100, []), RangeError);
  assert.throws(() => splitCents(100, [0, 0]), RangeError);
  assert.throws(() => splitCents(10.5, [1]), RangeError);
});
Notes

The weighted case passes on the starting code on purpose: rounding happens to work there, which is why the bug survived.

Sandbox run

Recorded by pnpm samples:verify. Our tests fail the build if this stops matching the task.

Checks passed5 tests pass on the reference solution; 3 of them fail on the starting code.
TestStarting codeReferenceSecond run
split.test.js::even split
passpasspasspass → pass
split.test.js::weighted split
passpasspasspass → pass
split.test.js::always adds up to the bill
Expected values to be strictly deep-equal: + actual - expected [ - 334, 333, 333, + 333 ]
failpasspassfail → pass
split.test.js::remainders go to the earliest people on ties
Expected values to be strictly deep-equal: + actual - expected [ 17, 17, 17, 17, + 17, + 17 - 16, - 16 ]
failpasspassfail → pass
split.test.js::rejects bad input
Missing expected exception (RangeError).
failpasspassfail → pass

JavaScript · Node.js 24 · node:test · canonset-sandbox-node:1 · 0.8 s · checked 2026-09-29 23:30 UTC

Output: Starting code (exit 1, 0.1 s)
✔ even split (1.4441ms)
✔ weighted split (0.164628ms)
✖ always adds up to the bill (1.162361ms)
✖ remainders go to the earliest people on ties (1.850176ms)
✖ rejects bad input (0.380776ms)
ℹ tests 5
ℹ suites 0
ℹ pass 2
ℹ fail 3
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 85.880885

✖ failing tests:

test at split.test.js:13:1
✖ always adds up to the bill (1.162361ms)
  AssertionError [ERR_ASSERTION]: Expected values to be strictly deep-equal:
  + actual - expected
  
    [
  -   334,
      333,
      333,
  +   333
    ]
  
      at TestContext.<anonymous> (file:///work/0-starter/split.test.js:14:10)
      at Test.runInAsyncScope (node:async_hooks:227:14)
      at Test.run (node:internal/test_runner/test:1402:25)
      at Test.processPendingSubtests (node:internal/test_runner/test:974:18)
      at Test.postRun (node:internal/test_runner/test:1542:19)
      at Test.run (node:internal/test_runner/test:1467:12)
      at async Test.processPendingSubtests (node:internal/test_runner/test:974:7) {
    generatedMessage: true,
    code: 'ERR_ASSERTION',
    actual: [ 333, 333, 333 ],
    expected: [ 334, 333, 333 ],
    operator: 'deepStrictEqual',
    diff: 'simple'
  }

test at split.test.js:17:1
✖ remainders go to the earliest people on ties (1.850176ms)
  AssertionError [ERR_ASSERTION]: Expected values to be strictly deep-equal:
  + actual - expected
  
    [
      17,
      17,
      17,
      17,
  +   17,
  +   17
  -   16,
  -   16
    ]
  
      at TestContext.<anonymous> (file:///work/0-starter/split.test.js:18:10)
      at Test.runInAsyncScope (node:async_hooks:227:14)
      at Test.run (node:internal/test_runner/test:1402:25)
      at Test.processPendingSubtests (node:internal/test_runner/test:974:18)
      at Test.postRun (node:internal/test_runner/test:1542:19)
      at Test.run (node:internal/test_runner/test:1467:12)
      at async Test.processPendingSubtests (node:internal/test_runner/test:974:7) {
    generatedMessage: true,
    code: 'ERR_ASSERTION',
    actual: [ 17, 17, 17, 17, 17, 17 ],
    expected: [ 17, 17, 17, 17, 16, 16 ],
    operator: 'deepStrictEqual',
    diff: 'simple'
  }

test at split.test.js:21:1
✖ rejects bad input (0.380776ms)
  AssertionError [ERR_ASSERTION]: Missing expected exception (RangeError).
      at TestContext.<anonymous> (file:///work/0-starter/split.test.js:22:10)
      at Test.runInAsyncScope (node:async_hooks:227:14)
      at Test.run (node:internal/test_runner/test:1402:25)
      at Test.processPendingSubtests (node:internal/test_runner/test:974:18)
      at Test.postRun (node:internal/test_runner/test:1542:19)
      at Test.run (node:internal/test_runner/test:1467:12)
      at async Test.processPendingSubtests (node:internal/test_runner/test:974:7) {
    generatedMessage: false,
    code: 'ERR_ASSERTION',
    actual: undefined,
    operator: 'throws',
    diff: 'simple'
  }
Output: Reference solution (exit 0, 0.2 s)
✔ even split (1.519129ms)
✔ weighted split (0.19241ms)
✔ always adds up to the bill (0.145201ms)
✔ remainders go to the earliest people on ties (0.994909ms)
✔ rejects bad input (0.513997ms)
ℹ tests 5
ℹ suites 0
ℹ pass 5
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 100.848422
Output: Reference, second run (exit 0, 0.1 s)
✔ even split (1.737696ms)
✔ weighted split (0.18922ms)
✔ always adds up to the bill (0.142223ms)
✔ remainders go to the earliest people on ties (1.189266ms)
✔ rejects bad input (0.520063ms)
ℹ tests 5
ℹ suites 0
ℹ pass 5
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 90.521145