AI Models · Tutorial 01
Compare AI models on a task you actually do
Test two models on the same small task. Compare factual errors, corrections, time, and cost before choosing a daily tool.

0 of 2 complete
Essentials · Step 1 of 2 · View the route
Last checked and updated: September 8, 2026. Editorial and source review; exercises below are authored practice, not reported benchmark runs.
You do not need a permanent ranking of every model. You need to know whether the one you are using can finish your work without creating another editing job.
In this lesson, compare two models on a short update. You will leave with a small evidence sheet and a reason for your choice. Use access you already have; buying a second subscription is not part of the exercise.
Start with a task you can judge
If you have not used an assistant before, do the 15-minute first exercise first. It shows how to spot invented details in a plausible answer.
Use these fictional notes for both models:
Monday: Interviewed 4 customers. 3 could not find the export button.
Tuesday: Moved the export button in a prototype. Nobody has tested it yet.
Wednesday: Sam offered to arrange a follow-up test. No date agreed.
Give both models the same brief in fresh conversations:
Write an update for a teammate who missed this week.
Use only the notes below. Keep it under 120 words.
Include what we learned, what changed, and what remains untested.
Do not invent a deadline, customer quote, or successful test.
Finish with one proposed next action, labelled as a proposal.
[Paste the three lines of notes here.]
Record the date, exact model names shown in the interface, and any settings you changed. If the two apps have different tools or hidden defaults, call this a comparison of those setups. You have not isolated the model alone.
Score the answer before reading its explanation
Write your checks before comparing the outputs. Otherwise, attractive phrasing can persuade you to accept an error.
| Check | Pass condition |
|---|---|
| Numbers | Four interviews and three reports remain accurate. |
| Evidence | The answer says the new position is untested. |
| Uncertainty | It does not invent a date or confirmed commitment from Sam. |
| Usefulness | The proposed next action follows from the notes. |
| Format | The update stays under 120 words. |
An answer must pass all five for this task. Do not average away a fabricated result because the writing sounds good.
For each failure, give one precise correction and record the time you spend. Save the original and corrected answers. The tool that produces the prettier first draft may still leave you with more work.
Keep a comparison sheet
Copy this table into your notes. The blank cells are for your observations; these are not published benchmark results.
| Observation | Setup A | Setup B |
|---|---|---|
| Model, version, app, date | ||
| Checks passed on first attempt, out of 5 | ||
| Corrections needed | ||
| Time waiting for answers | ||
| Time spent reviewing and correcting | ||
| Observed API cost or subscription usage indication | ||
| Final answer accepted? |
If an app does not show per-task cost, write “not available.” A subscription price is not a measured cost per answer, and a usage percentage is not a token count.
Now repeat the comparison with two other small sets of notes whose facts you can verify. Three examples can reveal a repeated failure, but they are too few to establish a general success rate. Keep the conclusion narrow: which setup worked better on these updates?
What a tier list can tell you
A tier list is an editor’s ranking under particular criteria. Before using it, look for the task, model version, tool access, settings, test date, and evidence behind the position. A coding result does not automatically predict good research or careful spreadsheet work.
Use our model benchmarks to find candidates and inspect examples. Then test a candidate against the job you care about. Model-family guides in this topic are reference reading, not prerequisites for this exercise.
Make a decision you can revisit
Write a short decision note:
For short updates from supplied notes, I will use [setup].
Across my three examples, it needed [observed corrections].
I chose it because [quality, review time, or observed cost].
I will retest if the model changes or it starts missing [specific check].
If neither setup passes, first check whether your sources and instructions are sufficient. An upgrade cannot supply a missing test date. If both pass, choose using the time, access, and cost you actually observed.
You are finished when someone else can read your sheet and understand the choice. Continue with hosted versus open-weight models if you also need to decide where the model should run.