Lab

How we test

Methodology

The Lab reviews models and AI tools under the same discipline the markets side applies to companies: real use, declared standing, boundaries in writing. This page is the current practice — what counts as a test here, what I'm allowed to claim, and what I won't trade for access.

Last reviewed

What gets tested

Real work, not synthetic benchmarks

I test models and AI tools by using them in real work — the research pipeline behind this site, the build projects in this lab, the code that ships here. A tool earns a verdict by doing an actual job with something at stake, not by clearing a synthetic benchmark suite.

That narrows coverage, and I accept the trade. A benchmark table can rank fifty models in an afternoon; a real-work test covers only what the work actually touches. If a tool hasn't entered the work, it doesn't get a verdict here — however loud the launch.

Hands-on before verdicts

The seat test

The standing rule: hands-on before verdicts. Every claim in a Lab piece is graded by the seat test — which seat was I in when I learned this?

  • Ran itI used the tool in real work, in my own seat. The only standing that supports a recommendation.
  • AdjacentI've worked the same stack or the same problem, but not this exact tool. I can explain mechanism and context; the piece says so, and the verdict waits.
  • Read itI've only read about it. Read-it-only standing never produces a recommendation — at most it produces coverage, labeled as coverage.

The standing is declared in the piece itself, so you can weigh every claim by the seat it was written from.

Disclosure

Labeled on the piece, not buried
  • Sponsored work is labeled. A paid or sponsored piece carries a visible Sponsored badge naming the sponsor — on the piece itself, not buried in a footer.
  • Partner links are disclosed. Some outbound links are tracked partner links; the pages that carry them say so. A commission never decides what gets covered or how anything ranks.
  • Positions are disclosed. On markets surfaces the published book is the disclosure — if I hold it, it's in the published positions. How that works is documented in the investing research methodology.

Corrections

In place, with a note

When something here turns out to be wrong, it gets corrected in place, with a note on the piece saying what changed. No silent rewrites, no quiet deletions. A wrong verdict is information; hiding it would be the real failure.

What I don't do

Boundaries
  • No pay-for-review. A verdict is not for sale, in any packaging. Sponsorship buys a labeled slot, never a conclusion.
  • No embargoed scores traded for access. If early access comes with conditions on what the verdict can be, I decline the access.
  • No vendor benchmark reprints presented as independent. A vendor's numbers are labeled as the vendor's. If I didn't run it, it is not presented as mine.

What this page will grow into

Intent, not claim

Stated as intent, not as a claim about today: as the testing corpus builds, this page grows into per-test result pages — the task, the setup, the standing, and the verdict, one page per test, linked from here. Until those pages exist, nothing on this page pretends they do; for now the record lives in the pieces themselves.

For the plain-English vocabulary behind the testing — inference, evals, agents — see the Explained library.