AI Models · Tutorial 01

Compare AI models on a task you actually do

Test two models on the same small task. Compare factual errors, corrections, time, and cost before choosing a daily tool.

A hand routes paper task parcels through four stepped levels of a mechanical model-tier tower.
Reading time
7 min
Last updated
September 2026

0 of 2 complete

Complete & next →

Essentials · Step 1 of 2 · View the route

Last checked and updated: September 8, 2026. Editorial and source review; exercises below are authored practice, not reported benchmark runs.

You do not need a permanent ranking of every model. You need to know whether the one you are using can finish your work without creating another editing job.

In this lesson, compare two models on a short update. You will leave with a small evidence sheet and a reason for your choice. Use access you already have; buying a second subscription is not part of the exercise.

Start with a task you can judge

If you have not used an assistant before, do the 15-minute first exercise first. It shows how to spot invented details in a plausible answer.

Use these fictional notes for both models:

Monday: Interviewed 4 customers. 3 could not find the export button.
Tuesday: Moved the export button in a prototype. Nobody has tested it yet.
Wednesday: Sam offered to arrange a follow-up test. No date agreed.

Give both models the same brief in fresh conversations:

Write an update for a teammate who missed this week.
Use only the notes below. Keep it under 120 words.
Include what we learned, what changed, and what remains untested.
Do not invent a deadline, customer quote, or successful test.
Finish with one proposed next action, labelled as a proposal.

[Paste the three lines of notes here.]

Record the date, exact model names shown in the interface, and any settings you changed. If the two apps have different tools or hidden defaults, call this a comparison of those setups. You have not isolated the model alone.

Score the answer before reading its explanation

Write your checks before comparing the outputs. Otherwise, attractive phrasing can persuade you to accept an error.

CheckPass condition
NumbersFour interviews and three reports remain accurate.
EvidenceThe answer says the new position is untested.
UncertaintyIt does not invent a date or confirmed commitment from Sam.
UsefulnessThe proposed next action follows from the notes.
FormatThe update stays under 120 words.

An answer must pass all five for this task. Do not average away a fabricated result because the writing sounds good.

For each failure, give one precise correction and record the time you spend. Save the original and corrected answers. The tool that produces the prettier first draft may still leave you with more work.

Keep a comparison sheet

Copy this table into your notes. The blank cells are for your observations; these are not published benchmark results.

ObservationSetup ASetup B
Model, version, app, date
Checks passed on first attempt, out of 5
Corrections needed
Time waiting for answers
Time spent reviewing and correcting
Observed API cost or subscription usage indication
Final answer accepted?

If an app does not show per-task cost, write “not available.” A subscription price is not a measured cost per answer, and a usage percentage is not a token count.

Now repeat the comparison with two other small sets of notes whose facts you can verify. Three examples can reveal a repeated failure, but they are too few to establish a general success rate. Keep the conclusion narrow: which setup worked better on these updates?

What a tier list can tell you

A tier list is an editor’s ranking under particular criteria. Before using it, look for the task, model version, tool access, settings, test date, and evidence behind the position. A coding result does not automatically predict good research or careful spreadsheet work.

Use our model benchmarks to find candidates and inspect examples. Then test a candidate against the job you care about. Model-family guides in this topic are reference reading, not prerequisites for this exercise.

Make a decision you can revisit

Write a short decision note:

For short updates from supplied notes, I will use [setup].
Across my three examples, it needed [observed corrections].
I chose it because [quality, review time, or observed cost].
I will retest if the model changes or it starts missing [specific check].

If neither setup passes, first check whether your sources and instructions are sufficient. An upgrade cannot supply a missing test date. If both pass, choose using the time, access, and cost you actually observed.

You are finished when someone else can read your sheet and understand the choice. Continue with hosted versus open-weight models if you also need to decide where the model should run.

Check your understanding

Q1.The update invents a Friday test date but reads well. How should you score it?
Q2.Two apps use different tools and default settings. What have you compared?
Q3.Your app does not show per-answer cost. What belongs in the cost cell?
Q4.One model passes three sample updates. What can you conclude?