Canonset
← Samples

Verified coding tasks

Tasks that prove their own tests

Each task is what a coding agent would receive: a problem statement and starting code, plus hidden tests and a reference solution. Our sandbox ran every one: the tests fail on the starting code, pass twice on the reference, and have no network access.

  • 12 tasks
  • Python 3.12, TypeScript, JavaScript
  • 33 fail→pass tests
  • Bug fixes, features, two security fixes

Make the LRU cache evict the least recently used entry

code-02 · what the agent sees is the problem statement and the starting code

mediumPython 3.12 · pytest
Problem statement

cache/lru.py has an LRUCache that is supposed to evict the least recently used key when it is full. In production it evicts keys that were read a moment ago: it evicts in insertion order, ignoring reads and updates. Fix it so that get and put both count as a use, eviction removes the least recently used key, and a capacity below 1 raises ValueError. Keep the public interface: LRUCache(capacity), get(key, default=None), put(key, value) and len().

Starting code
cache/lru.py
class LRUCache:
    def __init__(self, capacity: int):
        self.capacity = capacity
        self._data = {}

    def get(self, key, default=None):
        return self._data.get(key, default)

    def put(self, key, value):
        if key not in self._data and len(self._data) >= self.capacity:
            oldest = next(iter(self._data))
            del self._data[oldest]
        self._data[key] = value

    def __len__(self):
        return len(self._data)
Reference solution
cache/lru.py
from collections import OrderedDict


class LRUCache:
    def __init__(self, capacity: int):
        if capacity < 1:
            raise ValueError("capacity must be at least 1")
        self.capacity = capacity
        self._data: OrderedDict = OrderedDict()

    def get(self, key, default=None):
        if key not in self._data:
            return default
        self._data.move_to_end(key)
        return self._data[key]

    def put(self, key, value):
        if key in self._data:
            self._data.move_to_end(key)
        self._data[key] = value
        if len(self._data) > self.capacity:
            self._data.popitem(last=False)

    def __len__(self):
        return len(self._data)
Tests (hidden from the agent)
tests/test_lru.py
import pytest

from cache.lru import LRUCache


def test_get_and_put():
    cache = LRUCache(2)
    cache.put("a", 1)
    assert cache.get("a") == 1
    assert cache.get("missing", "default") == "default"


def test_capacity_is_respected():
    cache = LRUCache(3)
    for i in range(10):
        cache.put(i, i)
    assert len(cache) == 3


def test_read_counts_as_use():
    cache = LRUCache(2)
    cache.put("a", 1)
    cache.put("b", 2)
    cache.get("a")
    cache.put("c", 3)
    assert cache.get("a") == 1
    assert cache.get("b") is None


def test_update_counts_as_use():
    cache = LRUCache(2)
    cache.put("a", 1)
    cache.put("b", 2)
    cache.put("a", 10)
    cache.put("c", 3)
    assert cache.get("a") == 10
    assert cache.get("b") is None


def test_capacity_must_be_positive():
    with pytest.raises(ValueError):
        LRUCache(0)
Notes

Tests use only the public interface, so any correct implementation passes (OrderedDict or a linked list).

Sandbox run

Recorded by pnpm samples:verify. Our tests fail the build if this stops matching the task.

Checks passed5 tests pass on the reference solution; 3 of them fail on the starting code.
TestStarting codeReferenceSecond run
tests/test_lru.py::test_get_and_put
passpasspasspass → pass
tests/test_lru.py::test_capacity_is_respected
passpasspasspass → pass
tests/test_lru.py::test_read_counts_as_use
AssertionError: assert None == 1 + where None = get('a') + where get = <cache.lru.LRUCache object at 0x733ec6f38140>.get
failpasspassfail → pass
tests/test_lru.py::test_update_counts_as_use
AssertionError: assert None == 10 + where None = get('a') + where get = <cache.lru.LRUCache object at 0x733ec6f0fcb0>.get
failpasspassfail → pass
tests/test_lru.py::test_capacity_must_be_positive
Failed: DID NOT RAISE ValueError
failpasspassfail → pass

Python 3.12 · pytest · canonset-sandbox-python:1 · 1.3 s · checked 2026-09-29 23:30 UTC

Output: Starting code (exit 1, 0.3 s)
..FFF                                                                    [100%]
=================================== FAILURES ===================================
___________________________ test_read_counts_as_use ____________________________

    def test_read_counts_as_use():
        cache = LRUCache(2)
        cache.put("a", 1)
        cache.put("b", 2)
        cache.get("a")
        cache.put("c", 3)
>       assert cache.get("a") == 1
E       AssertionError: assert None == 1
E        +  where None = get('a')
E        +    where get = <cache.lru.LRUCache object at 0x733ec6f38140>.get

tests/test_lru.py:26: AssertionError
__________________________ test_update_counts_as_use ___________________________

    def test_update_counts_as_use():
        cache = LRUCache(2)
        cache.put("a", 1)
        cache.put("b", 2)
        cache.put("a", 10)
        cache.put("c", 3)
>       assert cache.get("a") == 10
E       AssertionError: assert None == 10
E        +  where None = get('a')
E        +    where get = <cache.lru.LRUCache object at 0x733ec6f0fcb0>.get

tests/test_lru.py:36: AssertionError
________________________ test_capacity_must_be_positive ________________________

    def test_capacity_must_be_positive():
>       with pytest.raises(ValueError):
             ^^^^^^^^^^^^^^^^^^^^^^^^^
E       Failed: DID NOT RAISE ValueError

tests/test_lru.py:41: Failed
=========================== short test summary info ============================
FAILED tests/test_lru.py::test_read_counts_as_use - AssertionError: assert No...
FAILED tests/test_lru.py::test_update_counts_as_use - AssertionError: assert ...
FAILED tests/test_lru.py::test_capacity_must_be_positive - Failed: DID NOT RA...
3 failed, 2 passed in 0.04s
Output: Reference solution (exit 0, 0.3 s)
.....                                                                    [100%]
5 passed in 0.02s
Output: Reference, second run (exit 0, 0.3 s)
.....                                                                    [100%]
5 passed in 0.02s