Keep going
Keep learning with Rundown Pro
Pro adds the full on-demand course library, certifications, group office hours, and eligible member perks.
$29/mo, cancel anytime
Guides published sep 30, 2026
In this guide, you will learn how to compare AI models on tasks you actually do and decide which subscriptions are worth paying for.
A simple spreadsheet of your recurring tasks, model scores, and the tools worth keeping. You'll also have saved prompts to use when a new model comes out.
Save the prompts you create in your spreadsheet, along with the model names, answers, and scores. When a new model comes out, run the same prompt and compare it with your current choice. You'll already have a test ready to go.
List tasks you do often. Start with an everyday task and a work task. For each one, write down how often you do it, what a good result looks like, and the tool you use now. Leave the tool blank if you're still choosing.
We started with dinner planning and a report for a boss. For dinner, we wanted an easy-to-follow recipe. For work, we wanted a short report with the right numbers and useful recommendations.
Track each subscription's monthly cost once, on a separate tab. That makes it easier to see what you're paying without counting the same plan for every task.
You can have AI build the sheet for you. Here's a prompt you can adapt:
Make me a spreadsheet for comparing AI models on tasks I actually do. Start with [your everyday task] and [your work task]. Include how often I do each task, what a good result looks like, and my current tool.
Give me a place to score each model for following the prompt, clarity, and tone, plus notes on formatting and anything I would need to fix. Add a column for the prompt so I can save it. Let me pick a primary tool and a backup for each task. Put my subscriptions and monthly costs on a separate tab.We changed ours into a matrix during the test so the answers were easier to compare. You can adjust yours as you go too.
Open OpenRouter's comparison page and select Flagship models. This gives you current top models from OpenAI, Anthropic, and Google, the companies behind ChatGPT, Claude, and Gemini.
Then open OpenRouter Chat, sign in, and select those models in the chat playground. This is where you send the same prompt to several models and read the answers side by side. You can add other models too. We tried Muse and Grok alongside the main three.
Pro tip: OpenRouter is the most cost-effective way we've found to test several flagship models. Our short test responses cost a few cents each. Budget about 25 cents for a short comparison; the total depends on the models and how much they write.
Add credits before testing paid models. These pay for your OpenRouter usage separately from any app subscriptions you have. Check model prices and your credit balance as you go; the OpenRouter FAQ explains how billing works.
Note the exact model names in your sheet so you know what you tested.
Think of a task you do often and write a prompt detailing it. Include the context, your limits, and how you want the answer written. Send that same prompt to the models, starting a fresh conversation for each task.
We tried dinner planning, a workshop debrief memo, and box office research. Here are versions you can try or adapt to your own tasks.
Start with dinner planning. Give it a time limit, available ingredients, and a word limit:
I need dinner for two in 45 minutes. I have spaghetti, a can of chickpeas, a can of chopped tomatoes, spinach, garlic, olive oil, salt, and pepper. I have one saucepan, one frying pan, and a stove, but no oven. I don’t want to shop or use other ingredients.
Choose one dinner and give me a timed plan. Put quantities inside the steps so I don’t have to keep scrolling. Include one way to avoid a bland or watery result. Explain it like a calm friend helping a tired person. Keep it under 300 words and don’t give me alternative meals.Read the steps and see whether you'd follow them. We didn't cook the recipes, so our dinner test was inconclusive. It still gave us a feel for the writing and how easy each answer was to use.
Next, try a work task where you know the facts. We asked for a workshop memo from notes with conflicting registration numbers, a capacity limit, a budget, and deadlines. Here's a sample prompt:
Turn these fictional project notes into a decision brief for my manager. Use only these facts:
- A customer workshop is planned for May 15; invitations should go out May 1.
- Sales says 80 customers are confirmed. The registration sheet has 52 unique registrations plus 28 duplicate rows.
- The venue holds 60 people, including our six staff. Walk-ins are not allowed.
- Budget cap: $12,000. Venue: $4,200. Production: $3,300. Travel: $1,800. Catering: $900.
- The venue deposit is refundable through April 25.
- Our speaker is tentative. Priya will confirm availability by April 24.
- Nate owns the invitation email and needs the final capacity and speaker decision by April 26.
Recommend proceed, pause, or change the plan. State the usable customer capacity, remaining customer places, and budget remaining. Identify the conflicting information without silently choosing the convenient version. Give a dated action list using only the named owners; mark any missing owner as unassigned. Finish with a manager-ready email of no more than 100 words. Keep the entire response under 350 words. Do not invent approvals, attendees, or bookings.For this example, the answer should allow 54 customers after reserving six places for staff. With 52 unique registrations, two customer places remain. Listed costs total $10,200, leaving $1,800 before unlisted costs. The action list should keep the April 24, 25, and 26 deadlines in order.
The first two tasks gave us similar answers, so we tried something harder: box office research. Here's a version you can use:
Find ten major studio movies releasing in [current year] or [next year] that you think will be important. Make a table with the title, release date, reported production budget, and your box-office prediction. Cite sources for the release dates and budgets. Mark unknown budgets as unknown, and clearly label your predictions. If a movie has already released, separate its actual box-office results from a forecast.
Open the sources to check the dates and budgets. In our test, Gemini included an already-released movie with a prediction. That's a useful mistake to note when you're comparing research answers.
Score each answer from 1 to 5 for following the prompt, clarity, and tone. Use 1 for poor, 3 for usable with edits, and 5 for an answer you'd use as it is.
Formatting is another major difference you'll notice. Does the answer put the recommendation first? Are the steps easy to follow? Would you rather read a table or several paragraphs? Put those preferences in your notes.
Claude's workshop answer started with a recommendation, put actions in a table, and included an email. We liked that layout. Claude and ChatGPT both appeared to follow the prompt, while Gemini's formatting and tone were less to our taste. You may prefer something different.
We also liked the dropdowns in Muse's research answer. Grok gave us a wall of text for that task, and its dinner answer didn't finish. Mark an incomplete answer as DNF and rerun it before treating that as a recurring problem.
Check the facts alongside the scores. A clear memo with the wrong budget still needs fixing. A short note like “easy to scan, but the budget is wrong” tells you more than a score alone. If speed matters to you, add that too.
Test five tasks, one a day. Use questions you actually ask and work you expect AI to help with. Pick a primary tool for each task and a backup if it helps with something your primary struggles with.
Then try those jobs in the app you'd keep. The experience in ChatGPT, Claude, or Gemini may differ from OpenRouter. Use the model available to your account, try your files and follow-up questions, and see whether it does what you need.
Look for subscriptions covering the same jobs. If one app handles both your everyday and work tasks, you may be able to cut an overlapping plan. If you're choosing your first subscription, use these tests to decide which app is worth paying for.
Keep going
Pro adds the full on-demand course library, certifications, group office hours, and eligible member perks.
$29/mo, cancel anytime