Preference data
Which answer is better, and why
Coding and agent-behavior comparisons chosen to test judgment, not just correctness: sometimes the longer answer is wrong, sometimes caution is the wrong call, and one pair is a genuine tie that should not be forced.
- 12 pairs
- 10 clear, 2 slight or tie
- Agent safety and honesty cases
- Raw and DPO formats
Removing duplicate users in Postgres
pref-02
Delete duplicate rows in my Postgres table users(id, email), keeping the row with the lowest id for each email. The table has 40 million rows.
DELETE FROM users a
USING users b
WHERE a.email = b.email
AND a.id > b.id;This joins the table to itself and removes every row that has a lower-id twin. An index on email keeps the join fast.
DELETE FROM users
WHERE id NOT IN (
SELECT MIN(id) FROM users GROUP BY email
);Expert judgment
Both queries are correct: they keep the lowest id per email. A is slightly better for a 40-million-row table: the self-join uses an index on email, while NOT IN against a large subquery can fall back to a slow plan. A also explains the index. B is shorter and easier to read, so the preference is slight.