> ## Content Index
> Fetch the complete content index at: https://theseguysknow.io/llms.txt
> Use this file to discover other available public pages before exploring further.

# How to Compare AI Models: The TGK Testing Method
- URL: https://theseguysknow.io/how-to-compare-ai-models/
- Published: 2026-09-09T07:23:19.000Z
- Updated: 2026-09-09T18:23:38.000Z
- Description: This is the method TGK uses to compare AI models without hiding rerolls, failed attempts, different settings or the limits of each test.
- Author: Mike Hazard
- Tags: AI & Tech, AI Tools & Models, #how we know

AI model comparisons become suspiciously easy when somebody enters one prompt, generates until every model returns something attractive, chooses a favourite from each pile and places the survivors into a neat ranking. The finished page may still be entertaining, but it cannot tell you how many weak attempts disappeared, whether one model received better instructions or whether four images from one service were treated as a single chance while another model was judged by the first thing it produced.

This is the method we use for TGK's AI comparisons, written openly enough that readers, journalists or anybody running their own test can reuse it. A useful comparison starts by deciding exactly what question the test should answer, because the right method changes depending on whether you care about first-try reliability, repeatable quality, the best result a skilled user can obtain or the amount of money required to reach something usable; once that question is clear, the job is less mysterious, because you keep the important conditions aligned, record every exception and make the conclusion no larger than the evidence.

## Quick answer

To compare AI models properly, give each model the same task, source material, success criteria and number of attempts; use equivalent settings where an exact match is impossible; record the model version, provider, test date, cost, failures and every material exception; then judge the outputs against criteria chosen before seeing the results.

Use first outputs to compare immediate reliability, repeated runs to compare consistency and model-specific prompt optimisation to compare maximum potential, because these are different tests and should never be quietly combined. A credible comparison does not require pretending that every interface is identical, but it does require telling the reader where the conditions differed and how that may have affected the result.

## Decide what the comparison is meant to prove

The easiest way to ruin a model comparison is to begin generating before deciding what winning means, because a beautiful result, an accurate result, a cheap result and a reliable result may come from four different models. NIST's AI Risk Management Framework makes the same point in more formal language: suitable measurements depend on the purpose, audience and needs of the evaluation, while the selected methods and anything that could not be measured should be documented.

For a practical editorial test, the question can be simple, but it still needs to be stated before the results arrive:

| The question                                      | A suitable test                                        | What the result can honestly show                                  |
| ------------------------------------------------- | ------------------------------------------------------ | ------------------------------------------------------------------ |
| Which model works best on the first try?          | One shared task and the first output from each model   | What a user received after one attempt under those conditions      |
| Which model is most consistent?                   | Several runs of the same tasks with unchanged settings | How often each model maintained its quality and followed the brief |
| Which model can produce the best finished result? | Disclosed rerolls, prompt changes and editing time     | The highest result reached with the stated amount of work          |
| Which model offers the best value?                | Record every paid attempt, failure and correction      | The money and time required to obtain a usable result              |

One output per model is enough for a first-output comparison, although it says very little about consistency and nothing sensible about a universal winner. Three runs can reveal obvious instability without turning a small website test into laboratory research, while stronger claims about reliability require a much larger sample and a method designed for statistics rather than an attractive comparison grid.

In our [five-model AI video comparison](https://theseguysknow.io/best-ai-video-generators-2026-tested/), keeping the first completed result made sense because instant understanding and the cost of a failed generation were part of the question. If we had been testing the highest possible quality after ten attempts and a full evening of editing, the same rule would have answered the wrong question.

## Keep the same task without forcing fake equality

Using the exact same prompt is often a good starting point, especially when several language models receive text through similar chat boxes, but identical wording can become artificial when one image or video service expects ordinary prose, another provides separate camera and negative-prompt controls, and a third does not support the requested duration or aspect ratio at all.

The stronger rule is to keep the meaning and difficulty of the task fixed, which means every model receives the same source files, required facts, subjects, actions, output goal and success criteria, while interface-specific formatting can change when necessary and every change is disclosed. A model should not receive a lovingly rewritten custom prompt simply because the reviewer likes it, yet forcing instructions into a format that the service does not understand would test the interface mismatch as much as the model.

Our [AI image generator test](https://theseguysknow.io/best-ai-image-generators/) followed this approach: every model received the same underlying briefs for a face, a hospital screen and a fragrance advertisement, while formatting changed slightly where an interface handled sections or shorter instructions better. When Midjourney returned four images in one batch, we used the first rather than quietly choosing the most flattering of four chances.

When a model cannot provide an equivalent feature, record the difference instead of hiding it. A video model without native audio should be marked as having no native audio; a service limited to 720p should not be described as matching a 2K result; and a blocked request should remain visible in the test record even though there is no output to score for quality.

## Record the model, provider and settings

A model name by itself is rarely enough to reproduce a result, because the same model can be served through its maker, a third-party platform or an API with different versions, default settings, compression, safety filters and prices. Our guide to [AI models, tools, platforms and providers](https://theseguysknow.io/ai-model-vs-tool-vs-platform-vs-provider/) explains the terminology, but the practical rule for comparisons is shorter: name the model and version, then name where and how you used it.

For every result, record the provider or platform, web or desktop interface, API route where relevant, generation mode, resolution, aspect ratio, duration, audio setting, temperature or seed when either is available, date of the test and the amount charged. Dates matter because models and interfaces change quickly, while prices shown before a generation can also differ from the final charge.

If two platforms both advertise the same model, that can justify a provider comparison, although it should not be presented as proof that the underlying model itself changed. You may be measuring the access route, exposed settings and billing system around it, which is useful information as long as the article names it correctly.

## Keep the first result and count the rerolls

Keeping the first result prevents a reviewer from rescuing a favourite model through endless retries, although “first result” needs a precise definition. We normally mean the first completed output, then separately report any errors, moderation blocks or jobs that never completed before it; those failures cannot be judged as images, videos or written answers, but they still affect whether somebody can use the service and whether they have paid for nothing.

Rerolls belong in the test when consistency or usable cost is being measured, yet they must remain visible. Record how many generations were attempted, which one was selected, whether the prompt changed, what the retries cost and whether failed jobs were refunded. Our [10-second AI video cost test](https://theseguysknow.io/how-much-does-ai-video-cost/) showed why this matters, because the advertised cost of one generation and the likely cost of reaching one usable clip were very different numbers.

If the winning image survived three secret rerolls while the losing image stayed on its first attempt, the page is showing portfolio choices rather than a comparison, however scientific the scorecard beside it may look.

## Choose the judging criteria before seeing the outputs

“Which one looks best?” can be a legitimate question when visual preference is the whole point, but it becomes a weak substitute for testing whether a model followed the instructions. Research behind GenAI-Bench found that image and video models could create photorealistic work while still struggling with attributes, relationships and more complicated prompt logic, which is why visual finish and prompt accuracy should usually receive separate judgments.

The criteria should match the job. A writing model can be judged on factual accuracy, instruction following, usefulness, source quality and clarity; an image model may need accurate text, natural anatomy, composition, reference fidelity and correct product scale; while an action-video test may care about character identity, order of movement, physical contact, location stability and synchronized sound. Speed and price should be recorded separately, because a cheap result does not become visually better through arithmetic and an expensive result does not become accurate through confidence.

Blind the model names while judging when the outputs allow it, randomise their order and use more than one human reviewer when the conclusion depends heavily on taste. Automated judges can help with a larger set, but they are not neutral referees by default: the MT-Bench and Chatbot Arena research documented position, verbosity and self-enhancement biases in model-based judging, so an AI judge's score should be checked rather than accepted as a small digital certificate of truth.

## Publish enough evidence for readers to disagree

Stanford's HELM evaluations publish prompt-level details and emphasise reproducible results, while NIST recommends documenting the test sets, methods, metrics and tools used during evaluation. A small publication does not need to imitate a research institution, but the same principle makes an editorial comparison far more useful: show enough evidence that another person can understand how the winner was chosen and where they might judge it differently.

This compact record can sit inside any AI model comparison:

> **Test question:** 
> **Task and required outcome:** 
> **Source files or shared input:** 
> **Models and exact versions:** 
> **Provider, platform or interface:** 
> **Test date and settings:** 
> **Attempts per model:** 
> **Failed, blocked or refunded runs:** 
> **Price charged and time taken:** 
> **Judging criteria:** 
> **Prompt or setting exceptions:** 
> **What the result cannot prove:**

Publishing the exact prompt and first outputs is better than summarising them, while source files, screenshots and downloadable data are worth adding when the test is large enough to justify them. The [TGK How We Know collection](https://theseguysknow.io/how-we-know/) already gathers tests where we show what we tried, what happened, what it cost and where the conclusion becomes less certain; this page gives those comparisons one shared method to reference.

You can see this method applied in the [TGK AI Video Generators 2026 dataset](https://huggingface.co/datasets/These-Guys-Know/tgk-ai-video-generators-2026?ref=theseguysknow.io), including the prompts, providers, settings, prices, first outputs and failed attempts.

## What an AI comparison can and cannot prove

A well-documented practical test can show which model performed best on the stated tasks, through the named provider, with the recorded settings and attempts, on the published date. It can reveal useful differences in instruction following, quality, stability, price or access, and it can help a reader decide which option deserves their own money.

It cannot prove that the winner is the best model for every task, every user or every future version, while one surprising failure does not prove that a model always fails in the same way. The smaller the sample, the narrower the claim should be, and any result that depends on personal taste should say so instead of turning an opinion into a decimal point.

## Final verdict

A credible AI model comparison defines the question before testing, keeps the meaningful conditions aligned, records versions and providers, preserves first outputs and failures, separates rerolls from first-try results and judges each model against criteria chosen in advance. Exact equality between different services is rarely possible, so transparency about every compromise matters more than pretending the interfaces behaved identically.

Readers do not need every comparison to resemble a university paper, but they should be able to see what was tested, what changed and how far the conclusion travels. When those details are missing, the ranking may still tell you which image the author liked; it just cannot tell you which model you should trust.

---

#### Sources

- [NIST AI Risk Management Framework Playbook: Measure](https://airc.nist.gov/airmf-resources/playbook/measure/?ref=theseguysknow.io)
- [Stanford Center for Research on Foundation Models: HELM Capabilities](https://crfm.stanford.edu/2025/03/20/helm-capabilities.html?ref=theseguysknow.io)
- [Zheng et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,” NeurIPS 2023](https://papers.nips.cc/paper%5Ffiles/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets%5Fand%5FBenchmarks.html?ref=theseguysknow.io)
- [Li et al., “GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation”](https://arxiv.org/abs/2406.13743?ref=theseguysknow.io)