Verified coding tasks
Tasks that prove their own tests
Each task is what a coding agent would receive: a problem statement and starting code, plus hidden tests and a reference solution. Our sandbox ran every one: the tests fail on the starting code, pass twice on the reference, and have no network access.
- 12 tasks
- Python 3.12, TypeScript, JavaScript
- 33 fail→pass tests
- Bug fixes, features, two security fixes
Order semantic versions by the spec
code-06 · what the agent sees is the problem statement and the starting code
versions/semver.py compares version strings for our update checker. It compares them as plain strings, so 1.10.0 sorts before 1.9.0, and pre-releases such as 1.0.0-beta are offered as newer than 1.0.0. Implement compare(a, b), returning -1, 0 or 1, following Semantic Versioning 2.0.0 precedence: - compare major, minor and patch numerically; - a pre-release (1.0.0-alpha) is lower than the same version without one; - pre-release identifiers are compared left to right: numeric ones numerically, others as ASCII text, numeric lower than text, and a longer list wins when all earlier identifiers are equal; - build metadata after "+" is ignored.
def compare(a: str, b: str) -> int:
return (a > b) - (a < b)def _parse(version: str):
version = version.split("+", 1)[0]
core, _, pre = version.partition("-")
major, minor, patch = (int(part) for part in core.split("."))
return (major, minor, patch), pre.split(".") if pre else []
def _compare_identifiers(a: list[str], b: list[str]) -> int:
for x, y in zip(a, b):
if x == y:
continue
x_num, y_num = x.isdigit(), y.isdigit()
if x_num and y_num:
return (int(x) > int(y)) - (int(x) < int(y))
if x_num:
return -1
if y_num:
return 1
return (x > y) - (x < y)
return (len(a) > len(b)) - (len(a) < len(b))
def compare(a: str, b: str) -> int:
(core_a, pre_a), (core_b, pre_b) = _parse(a), _parse(b)
if core_a != core_b:
return (core_a > core_b) - (core_a < core_b)
if not pre_a and not pre_b:
return 0
if not pre_a:
return 1
if not pre_b:
return -1
return _compare_identifiers(pre_a, pre_b)from versions.semver import compare
def test_equal():
assert compare("1.2.3", "1.2.3") == 0
def test_major_versions():
assert compare("1.0.0", "2.0.0") == -1
def test_numeric_not_lexical():
assert compare("1.10.0", "1.9.0") == 1
def test_prerelease_is_lower_than_release():
assert compare("1.0.0-alpha", "1.0.0") == -1
def test_prerelease_chain_from_the_spec():
chain = ["1.0.0-alpha", "1.0.0-alpha.1", "1.0.0-alpha.beta", "1.0.0-beta", "1.0.0-beta.2", "1.0.0-beta.11", "1.0.0-rc.1", "1.0.0"]
for lower, higher in zip(chain, chain[1:]):
assert compare(lower, higher) == -1, (lower, higher)
assert compare(higher, lower) == 1, (higher, lower)
def test_build_metadata_is_ignored():
assert compare("1.0.0+build.5", "1.0.0+build.9") == 0The chain test is the example ordering from the SemVer 2.0.0 specification, checked in both directions.
Sandbox run
Recorded by pnpm samples:verify. Our tests fail the build if this stops matching the task.
| Test | Starting code | Reference | Second run | |
|---|---|---|---|---|
tests/test_semver.py::test_equal | pass | pass | pass | pass → pass |
tests/test_semver.py::test_major_versions | pass | pass | pass | pass → pass |
tests/test_semver.py::test_numeric_not_lexical AssertionError: assert -1 == 1
+ where -1 = compare('1.10.0', '1.9.0') | fail | pass | pass | fail → pass |
tests/test_semver.py::test_prerelease_is_lower_than_release AssertionError: assert 1 == -1
+ where 1 = compare('1.0.0-alpha', '1.0.0') | fail | pass | pass | fail → pass |
tests/test_semver.py::test_prerelease_chain_from_the_spec AssertionError: ('1.0.0-beta.2', '1.0.0-beta.11')
assert 1 == -1
+ where 1 = compare('1.0.0-beta.2', '1.0.0-beta.11') | fail | pass | pass | fail → pass |
tests/test_semver.py::test_build_metadata_is_ignored AssertionError: assert -1 == 0
+ where -1 = compare('1.0.0+build.5', '1.0.0+build.9') | fail | pass | pass | fail → pass |
Python 3.12 · pytest · canonset-sandbox-python:1 · 1.2 s · checked 2026-09-29 23:30 UTC
Output: Starting code (exit 1, 0.3 s)
..FFFF [100%]
=================================== FAILURES ===================================
___________________________ test_numeric_not_lexical ___________________________
def test_numeric_not_lexical():
> assert compare("1.10.0", "1.9.0") == 1
E AssertionError: assert -1 == 1
E + where -1 = compare('1.10.0', '1.9.0')
tests/test_semver.py:13: AssertionError
____________________ test_prerelease_is_lower_than_release _____________________
def test_prerelease_is_lower_than_release():
> assert compare("1.0.0-alpha", "1.0.0") == -1
E AssertionError: assert 1 == -1
E + where 1 = compare('1.0.0-alpha', '1.0.0')
tests/test_semver.py:17: AssertionError
_____________________ test_prerelease_chain_from_the_spec ______________________
def test_prerelease_chain_from_the_spec():
chain = ["1.0.0-alpha", "1.0.0-alpha.1", "1.0.0-alpha.beta", "1.0.0-beta", "1.0.0-beta.2", "1.0.0-beta.11", "1.0.0-rc.1", "1.0.0"]
for lower, higher in zip(chain, chain[1:]):
> assert compare(lower, higher) == -1, (lower, higher)
E AssertionError: ('1.0.0-beta.2', '1.0.0-beta.11')
E assert 1 == -1
E + where 1 = compare('1.0.0-beta.2', '1.0.0-beta.11')
tests/test_semver.py:23: AssertionError
________________________ test_build_metadata_is_ignored ________________________
def test_build_metadata_is_ignored():
> assert compare("1.0.0+build.5", "1.0.0+build.9") == 0
E AssertionError: assert -1 == 0
E + where -1 = compare('1.0.0+build.5', '1.0.0+build.9')
tests/test_semver.py:28: AssertionError
=========================== short test summary info ============================
FAILED tests/test_semver.py::test_numeric_not_lexical - AssertionError: asser...
FAILED tests/test_semver.py::test_prerelease_is_lower_than_release - Assertio...
FAILED tests/test_semver.py::test_prerelease_chain_from_the_spec - AssertionE...
FAILED tests/test_semver.py::test_build_metadata_is_ignored - AssertionError:...
4 failed, 2 passed in 0.04s
Output: Reference solution (exit 0, 0.3 s)
...... [100%] 6 passed in 0.02s
Output: Reference, second run (exit 0, 0.3 s)
...... [100%] 6 passed in 0.02s