New this week: 6 head-to-head tests added, including Claude vs Gemini for long documents.

See what changed →
Methodology

One brief, forty tasks, every tool in the category

A score on this site only means something if it was produced the same way for every contender. Here is the exact process we run before a hub goes live, and the rules that decide the order.

40 tasks per category6 scoring criteriaRe-scored every quarter
The process

From shortlist to published hub

Nothing is published from a vendor demo, a press release or a competitor round-up. Every line on a hub comes from a seat we paid for.

STEP 1

Define the job

We write the category brief first: who buys this, what they actually produce with it, and what a good week of output looks like. The brief is fixed before we look at any tool.

STEP 2

Build the shortlist

Anything with meaningful usage, a public price and an active product goes on the list. Abandoned tools and vapourware are cut here rather than ranked badly later.

STEP 3

Buy the seats

We pay full price on the plan a normal buyer would choose. That means we see the real limits, the real support response and the real invoice, including any minimum seat count.

STEP 4

Run the 40 tasks

Every tool gets the identical brief in the same order, with the same source material. Outputs are saved so two people can score them without knowing which tool produced what.

STEP 5

Score and price

Six criteria are scored independently, then we calculate the cost of a real working month at one, five and twenty seats so the cheap-looking option cannot hide behind a starter tier.

STEP 6

Publish and re-check

The hub goes live with a last-updated date. Prices and limits are re-checked weekly, and the full brief is re-run at least once a quarter or whenever a major release lands.

Scoring

The six things we score, and why

Output quality

Blind-scored against the brief. We care about how much editing the result needs before a client or a manager would accept it, not how impressive the first paragraph looks.

Real monthly cost

The invoice for a working month at the seat count you would actually buy, including minimums, overage and anything that is annual-only in disguise.

Speed under load

Time to a usable result on a normal weekday, not a benchmark at 3am. Queues and rate limits count, because they are what you feel on a deadline.

Data handling

Where content is stored, whether it is used for training by default, what the retention window is, and how clearly the vendor documents all three.

Export and lock-in

How much of your work leaves in a usable format if you cancel tomorrow. Proprietary-only export is a real cost and it is scored as one.

Support and stability

Response time on a genuine ticket, plus how often the product broke, changed pricing or removed a feature during the test window.

Ranking rules

How ties and edge cases are resolved

When two tools land within a point of each other, the cheaper real monthly cost wins. If cost is also level, the tool with the better export story wins, because that is the decision you cannot easily reverse later.

A tool can be removed from a hub between test cycles if it shuts down, changes pricing so severely that the review misleads, or ships a change that breaks the job it was recommended for. Removals are noted in the changelog rather than quietly deleted.

We do not run a single overall winner for every reader. Hubs are written around use cases, because the best tool for a solo freelancer on fifteen pounds a month is rarely the best tool for a team of twenty.

Conflicts of interest

What could bias us, and what we do about it

Scores are locked before any commercial link is added to a page. The person who scores a category does not handle affiliate relationships, and no vendor sees a review before it publishes.

If we have any other relationship with a company — a former client, an investor connection, anything at all — it is stated on the page. Read the full editorial standards and affiliate disclosure.

Scores locked pre-monetisationNo vendor previewsPublic changelogs
Questions

About the method

Why forty tasks?
Fewer than about thirty and one lucky output distorts the result. Many more than forty and we cannot afford to re-run the whole brief every quarter across every category. Forty is the point where the ranking stopped moving when we added more.
Are the tests blind?
Output scoring is. Results are stripped of branding and scored by two people independently, and we only reconcile the numbers afterwards. Pricing and data-handling checks are not blind because they require reading the vendor documentation.
How often is a hub updated?
Prices and limits weekly, the full brief at least quarterly, and immediately when a tool ships a major release. Every hub shows its last-updated date at the top.
Can a vendor challenge a score?
Yes, with evidence. If a test was run on a stale version or we misread a plan, we re-run it and correct the page. What we will not do is remove a fair criticism because it is unflattering.

See the method applied

Every hub on the site was built with this exact process. Start with the tool you already pay for.