Canonset
← Samples

Verified coding tasks

Tasks that prove their own tests

Each task is what a coding agent would receive: a problem statement and starting code, plus hidden tests and a reference solution. Our sandbox ran every one: the tests fail on the starting code, pass twice on the reference, and have no network access.

  • 12 tasks
  • Python 3.12, TypeScript, JavaScript
  • 33 fail→pass tests
  • Bug fixes, features, two security fixes

Order semantic versions by the spec

code-06 · what the agent sees is the problem statement and the starting code

hardPython 3.12 · pytest
Problem statement

versions/semver.py compares version strings for our update checker. It compares them as plain strings, so 1.10.0 sorts before 1.9.0, and pre-releases such as 1.0.0-beta are offered as newer than 1.0.0. Implement compare(a, b), returning -1, 0 or 1, following Semantic Versioning 2.0.0 precedence: - compare major, minor and patch numerically; - a pre-release (1.0.0-alpha) is lower than the same version without one; - pre-release identifiers are compared left to right: numeric ones numerically, others as ASCII text, numeric lower than text, and a longer list wins when all earlier identifiers are equal; - build metadata after "+" is ignored.

Starting code
versions/semver.py
def compare(a: str, b: str) -> int:
    return (a > b) - (a < b)
Reference solution
versions/semver.py
def _parse(version: str):
    version = version.split("+", 1)[0]
    core, _, pre = version.partition("-")
    major, minor, patch = (int(part) for part in core.split("."))
    return (major, minor, patch), pre.split(".") if pre else []


def _compare_identifiers(a: list[str], b: list[str]) -> int:
    for x, y in zip(a, b):
        if x == y:
            continue
        x_num, y_num = x.isdigit(), y.isdigit()
        if x_num and y_num:
            return (int(x) > int(y)) - (int(x) < int(y))
        if x_num:
            return -1
        if y_num:
            return 1
        return (x > y) - (x < y)
    return (len(a) > len(b)) - (len(a) < len(b))


def compare(a: str, b: str) -> int:
    (core_a, pre_a), (core_b, pre_b) = _parse(a), _parse(b)
    if core_a != core_b:
        return (core_a > core_b) - (core_a < core_b)
    if not pre_a and not pre_b:
        return 0
    if not pre_a:
        return 1
    if not pre_b:
        return -1
    return _compare_identifiers(pre_a, pre_b)
Tests (hidden from the agent)
tests/test_semver.py
from versions.semver import compare


def test_equal():
    assert compare("1.2.3", "1.2.3") == 0


def test_major_versions():
    assert compare("1.0.0", "2.0.0") == -1


def test_numeric_not_lexical():
    assert compare("1.10.0", "1.9.0") == 1


def test_prerelease_is_lower_than_release():
    assert compare("1.0.0-alpha", "1.0.0") == -1


def test_prerelease_chain_from_the_spec():
    chain = ["1.0.0-alpha", "1.0.0-alpha.1", "1.0.0-alpha.beta", "1.0.0-beta", "1.0.0-beta.2", "1.0.0-beta.11", "1.0.0-rc.1", "1.0.0"]
    for lower, higher in zip(chain, chain[1:]):
        assert compare(lower, higher) == -1, (lower, higher)
        assert compare(higher, lower) == 1, (higher, lower)


def test_build_metadata_is_ignored():
    assert compare("1.0.0+build.5", "1.0.0+build.9") == 0
Notes

The chain test is the example ordering from the SemVer 2.0.0 specification, checked in both directions.

Sandbox run

Recorded by pnpm samples:verify. Our tests fail the build if this stops matching the task.

Checks passed6 tests pass on the reference solution; 4 of them fail on the starting code.
TestStarting codeReferenceSecond run
tests/test_semver.py::test_equal
passpasspasspass → pass
tests/test_semver.py::test_major_versions
passpasspasspass → pass
tests/test_semver.py::test_numeric_not_lexical
AssertionError: assert -1 == 1 + where -1 = compare('1.10.0', '1.9.0')
failpasspassfail → pass
tests/test_semver.py::test_prerelease_is_lower_than_release
AssertionError: assert 1 == -1 + where 1 = compare('1.0.0-alpha', '1.0.0')
failpasspassfail → pass
tests/test_semver.py::test_prerelease_chain_from_the_spec
AssertionError: ('1.0.0-beta.2', '1.0.0-beta.11') assert 1 == -1 + where 1 = compare('1.0.0-beta.2', '1.0.0-beta.11')
failpasspassfail → pass
tests/test_semver.py::test_build_metadata_is_ignored
AssertionError: assert -1 == 0 + where -1 = compare('1.0.0+build.5', '1.0.0+build.9')
failpasspassfail → pass

Python 3.12 · pytest · canonset-sandbox-python:1 · 1.2 s · checked 2026-09-29 23:30 UTC

Output: Starting code (exit 1, 0.3 s)
..FFFF                                                                   [100%]
=================================== FAILURES ===================================
___________________________ test_numeric_not_lexical ___________________________

    def test_numeric_not_lexical():
>       assert compare("1.10.0", "1.9.0") == 1
E       AssertionError: assert -1 == 1
E        +  where -1 = compare('1.10.0', '1.9.0')

tests/test_semver.py:13: AssertionError
____________________ test_prerelease_is_lower_than_release _____________________

    def test_prerelease_is_lower_than_release():
>       assert compare("1.0.0-alpha", "1.0.0") == -1
E       AssertionError: assert 1 == -1
E        +  where 1 = compare('1.0.0-alpha', '1.0.0')

tests/test_semver.py:17: AssertionError
_____________________ test_prerelease_chain_from_the_spec ______________________

    def test_prerelease_chain_from_the_spec():
        chain = ["1.0.0-alpha", "1.0.0-alpha.1", "1.0.0-alpha.beta", "1.0.0-beta", "1.0.0-beta.2", "1.0.0-beta.11", "1.0.0-rc.1", "1.0.0"]
        for lower, higher in zip(chain, chain[1:]):
>           assert compare(lower, higher) == -1, (lower, higher)
E           AssertionError: ('1.0.0-beta.2', '1.0.0-beta.11')
E           assert 1 == -1
E            +  where 1 = compare('1.0.0-beta.2', '1.0.0-beta.11')

tests/test_semver.py:23: AssertionError
________________________ test_build_metadata_is_ignored ________________________

    def test_build_metadata_is_ignored():
>       assert compare("1.0.0+build.5", "1.0.0+build.9") == 0
E       AssertionError: assert -1 == 0
E        +  where -1 = compare('1.0.0+build.5', '1.0.0+build.9')

tests/test_semver.py:28: AssertionError
=========================== short test summary info ============================
FAILED tests/test_semver.py::test_numeric_not_lexical - AssertionError: asser...
FAILED tests/test_semver.py::test_prerelease_is_lower_than_release - Assertio...
FAILED tests/test_semver.py::test_prerelease_chain_from_the_spec - AssertionE...
FAILED tests/test_semver.py::test_build_metadata_is_ignored - AssertionError:...
4 failed, 2 passed in 0.04s
Output: Reference solution (exit 0, 0.3 s)
......                                                                   [100%]
6 passed in 0.02s
Output: Reference, second run (exit 0, 0.3 s)
......                                                                   [100%]
6 passed in 0.02s