Canonset
← Samples

Preference data

Which answer is better, and why

Coding and agent-behavior comparisons chosen to test judgment, not just correctness: sometimes the longer answer is wrong, sometimes caution is the wrong call, and one pair is a genuine tie that should not be forced.

  • 12 pairs
  • 10 clear, 2 slight or tie
  • Agent safety and honesty cases
  • Raw and DPO formats

Fast membership checks

pref-07

Prompt

I check whether values are in a list of 1 million strings, about 10,000 times per second. It's slow. What should I do?

Response A

Convert the list to a set once: allowed = set(values). Then x in allowed is O(1) on average instead of scanning the whole list each time.

Response B

Sort the list once, then use binary search:

import bisect
values.sort()
def contains(x):
    i = bisect.bisect_left(values, x)
    return i < len(values) and values[i] == x

Each lookup is O(log n).

Expert judgment

Prefers response AClear preference

Both fix the O(n) scan, but a set is the idiomatic answer: one line, faster (O(1) average versus about 20 comparisons per lookup) and with no helper to get wrong. B is correct but more code for a slower result.