Canonset
← Samples

Preference data

Which answer is better, and why

Coding and agent-behavior comparisons chosen to test judgment, not just correctness: sometimes the longer answer is wrong, sometimes caution is the wrong call, and one pair is a genuine tie that should not be forced.

  • 12 pairs
  • 10 clear, 2 slight or tie
  • Agent safety and honesty cases
  • Raw and DPO formats

Removing duplicate users in Postgres

pref-02

Prompt

Delete duplicate rows in my Postgres table users(id, email), keeping the row with the lowest id for each email. The table has 40 million rows.

Response A
DELETE FROM users a
USING users b
WHERE a.email = b.email
  AND a.id > b.id;

This joins the table to itself and removes every row that has a lower-id twin. An index on email keeps the join fast.

Response B
DELETE FROM users
WHERE id NOT IN (
  SELECT MIN(id) FROM users GROUP BY email
);

Expert judgment

Prefers response ASlight preference

Both queries are correct: they keep the lowest id per email. A is slightly better for a 40-million-row table: the self-join uses an index on email, while NOT IN against a large subquery can fall back to a slow plan. A also explains the index. B is shorter and easier to read, so the preference is slight.