Method

Every finding on this site is produced the same way. This page describes that process so you can judge the work rather than take it on trust — and so you can see exactly where my judgement enters, which is the part most worth arguing with.

How a survey narrows from papers to a published finding Seventy-four collected papers, shown as a grid of dots, reduce across four stages to six clusters, three judged practical, two repositories opened, and one published finding. 746321 paperscollected clustersidentified judgedpractical reposopened findingpublished
The shape of a typical survey. Most of the work is deciding what does not warrant a closer look — and saying so.

1. Defining the domain

A survey is only useful if the boundary is narrow enough to read exhaustively. Before collecting anything, I write down what is in scope and what is out, and I publish that definition with the finding.

The definition fixes:

  • The technical question the domain is organised around
  • The time window, usually the preceding eighteen to thirty months
  • Which venues and preprint servers I searched, and with what terms
  • The date collection closed

Preprints are included. A great deal of current AI work never reaches a reviewed venue, and excluding it would give a distorted picture. Where a preprint has since been withdrawn or substantially revised, I say so.

Work I could not access — paywalled, or published in a language I cannot read — is listed as a known gap rather than quietly omitted.

2. Classification

I sort proposals by the mechanism they actually use, not by how they are framed in the abstract. Papers routinely present a familiar technique as novel, or a narrow result as a general one, and grouping by claimed contribution produces a map that reflects marketing rather than substance.

The resulting taxonomy is mine. Someone else reading the same papers would group them differently, and I explain the reasoning behind each grouping so a reader who disagrees can see precisely where we diverge. Where a proposal sits awkwardly between groups, I say that rather than forcing it.

3. Assessing practical usability

For each group, one question: what would it take to use this outside a benchmark?

I look at:

  • Data. What is required, at what scale, and whether it is obtainable outside the lab that produced the paper.
  • Compute. What training and inference cost, in terms a reader can map onto hardware they might actually have.
  • Robustness. How the approach behaves when its assumptions are not met, and whether the paper tests that at all.
  • Maintenance. Dependencies, brittleness, and what breaks when the surrounding ecosystem moves.
  • Licensing. Whether the code and any released weights permit the uses a reader is likely to have in mind.

This assessment is a judgement, presented as a judgement. It is the least certain part of any finding and the part most likely to be wrong. I state my confidence and my reasons, and I would rather be visibly uncertain than falsely crisp.

4. Where it warrants it, running the code

A small number of proposals justify going further. I open a repository when the approach looks practically relevant, an implementation is public, and the result would materially change what a reader should do.

When I do, I record:

  • The exact commit, by hash, and the date I ran it
  • Hardware, operating system, and dependency versions
  • What I ran, including seeds and any deviation from the documented procedure
  • What completed, what did not, and where I stopped
  • How long I spent before stopping

That last point matters. Failing to get something running is often a statement about my time, patience, or setup rather than about the code. I distinguish between this did not work and I could not make this work, because they are different claims and only one of them is about the software.

Findings tied to a commit are true of that commit. Later versions may behave differently, and a repository that failed once may work perfectly a month later.

5. Publication and correction

Findings are written in my own words. Papers and repositories are linked to their sources rather than reproduced; where a short quotation is genuinely necessary to make a point precise, it is brief and attributed. I do not republish figures, tables, or code from the works I discuss.

Every finding carries the date it was published and the date of any subsequent revision. Corrections appear alongside the original text with the change noted, never silently folded in. If a correction is substantial enough to change the conclusion, it goes at the top.

Corrections and right of reply

If you are an author and believe I have got something wrong, write to contact@currybench.org. I will look again, and if I am wrong I will publish the correction with the same prominence as the original. If we continue to disagree, I will publish your response alongside my finding so a reader can weigh both.

I aim to respond within [ten working days].

Conflicts of interest

I have no affiliation with the research groups, companies, or institutions whose work I read, and I receive no payment from them. If that ever ceases to be true for a particular finding, the relationship will be declared on that finding itself, at the top, before the analysis.

This site is non-commercial. It sells nothing, carries no advertising, and accepts no sponsorship.