McCloskey.ai · ResourcesAll resources
← Resources The Leverage Ladder

Beyond the single prompt

Almost everyone stops at one prompt. Each level up is mostly just a few extra sentences. Here is the whole ladder.

An ascending staircase of techniques: a good prompt, an eval check, deep research and thinking, a loop, slash goal, and ultracode, rising from a single prompt to a team of agents. A team of agents Single prompt Ultracode a whole team of agents /goal a loop with a finish line A loop it fixes its own work Deep research and thinking An eval check it grades itself A good prompt one clear ask

Each level adds a little, and it is mostly just a few extra sentences. The eval check is it grades its own work, the loop is it fixes what it grades, ultracode is a whole team. The trick is not climbing to the top.

Climb to the lowest level that does the job, not the highest. Most work only needs the first two.

Not sure which level? Ask it to talk you down · the golden rule as a prompt
  • The cheapest answer that nails it beats the fanciest one that makes a mess and burns your plan.
  • Before you climb, ask: would the level below have done this? If yes, drop down.
Pick the right level
Here's a task I want to do with AI: [describe it].

Before we start, tell me the LOWEST-effort way to get this done well:
- Would a plain chat with good context do it? If so, stop there.
- If not, what's the cheapest step up that fits: a saved skill, deep research, deep thinking, sub-agents, a loop, /goal, or a scheduled task?
- Name anything I might over-reach on, and talk me down to the simpler option if it would do the job.

I want the cheapest thing that actually solves it, not the most impressive.
Where these levels live · Chat, Cowork, Claude Code
  • Claude chat (the app you already use): asking with context, Projects, deep research and deep thinking.
  • Cowork (a tab in the same app): your hands-off employee, where skills, connectors and scheduled tasks live.
  • Claude Code (the power-user tool): /goal, workflows and ultracode, which most professionals never open.

One honest note: making images and polished visual designs is a separate tool (Claude Design). Worth knowing, but not a rung on this ladder.

Level 1 A good prompt

What it is

One clear ask, with the context, the audience and the format you want spelled out.

Reach for it when

You want a single answer, draft or explanation right now.

Cost lowest
Copy, paste, fill the brackets
Help me with this clearly and in plain English.

What I need: [THE TASK OR QUESTION]
Audience: [WHO THIS IS FOR]
Format I want back: [BULLETS / SHORT EMAIL / STEP-BY-STEP]

Rules:
- Only use what I give you. If something is missing, ask one question first.
- Be concrete. No filler, no disclaimers.
- Return just the answer, nothing else.

Details:
[PASTE YOUR CONTEXT HERE]

New to prompting? How to Prompt breaks a good ask into three simple steps.

Level 2 An eval check

What it is

A good prompt that grades its own answer against a standard, fixes the weakest part once, then shows you only the improved version.

Reach for it when

You want one answer but better, with the obvious flaws already caught. The first taste of the AI checking itself.

Cost very low
Copy, paste, fill the brackets
Answer this, then check your own work once before you show me anything.

What I need: [THE TASK OR QUESTION]
What "good" means here: [THE STANDARD, AS A FEW CLEAR RULES, e.g. under 150 words, plain English, names one next step]

Do this:
1. Write your answer.
2. Rate it 1 to 10 against the rules above, and say in one line why.
3. Fix the single weakest part once.
4. Show me only the improved answer, then the score it earned.

Be a hard marker. Do not show me the first draft.
Side note · saving any level as a skill

Like a prompt and want it on tap? Tell the AI to save it as a skill, so you run it by name instead of retyping it. This works for every level on this page, not just this one. It is a convenience, not a step up the ladder. Just add: "save this as a reusable skill called [short-name] so I can run it again by name."

Go deeper: skills, properly · the side note, in full
  • A skill is a saved recipe: you teach it a task once, in plain words, and reuse it forever instead of re-explaining.
  • It captures your way of doing something: the steps, the standard a good result must hit, and the "we don't do it like that" rules.
  • The AI decides when to fire a skill from its description, so keep the description boring and exact, not clever.
  • A skill saves the task; a Project saves your context. Use both: the Project knows your world, the skill knows the job.
  • The trigger: the moment you find yourself writing the same kind of prompt a third time.
At work: a marketer saves a "LinkedIn post from a report" skill: take a link, pull the three best points, write three short posts in my voice, no hashtags. Next week it is one line: "run my LinkedIn-post skill on this link."
Build a reusable skill
Turn this into a reusable skill I can run again later.
Name it something obvious. Write it as: when to use it, the exact steps, and the standard a good result must hit.
The task: [describe the repeatable job in plain English, including your rules and your style]
Save it, then show me the one-line command I'll use to run it next time.

Level 3 Deep research and thinking

What it is

You tell it to slow down, gather sources or reason hard before it answers, instead of replying off the top of its head.

Reach for it when

Being right matters more than being fast, or the obvious answer feels too easy.

Cost medium
Copy, paste, fill the brackets
Take your time on this. Think it through before you answer.

The question: [THE DECISION OR TOPIC]
Why it matters: [WHAT THIS FEEDS INTO]

Do this:
1. Restate the real question in one sentence.
2. Lay out the 2 to 4 most credible answers, each with its strongest case.
3. For factual claims, name the source. Separate solid from uncertain.
4. Give me your recommendation and why, then what would change your mind.

If you are unsure, say so plainly. Do not invent figures or quotes.

For what genuine research depth looks like, see the studies.

When the stakes are high · two richer prompts

The prompt above is the everyday version. When a misstep would be expensive, split the job in two: research gathers the facts, thinking makes the call.

Deep research · a cited report you can trust
Do deep research and give me a cited report I can trust.

QUESTION: [the specific thing you need answered]
WHY IT MATTERS: [the decision this feeds, so you can judge what's relevant]
MUST COVER: [3-5 things the report has to address, e.g. real pricing, recent reviews, market size, the catch]
SOURCE BAR: prefer recent, primary, reputable sources. Flag anything you could only find in one place as "single-source, unverified."

Deliver:
1. A direct answer up front, one paragraph.
2. The evidence, grouped by point, each claim with a citation.
3. Where sources disagreed, and which you trust more and why.
4. What you could NOT confirm.

Don't pad it. If a section has no solid evidence, say so.
Deep thinking · attack it, then try to break it
Think hard about this before answering. Do NOT just agree with me.

DECISION: [the question or choice you're facing]
WHAT I'M LEANING TOWARD: [your current gut call, if any]
WHAT'S AT STAKE: [time, money, headcount, or reputation on the line]

Process:
1. Attack this from at least three independent angles (e.g. cost, risk, second-order effects, what a sceptic on the board would say). Treat them as separate viewpoints, not one.
2. Form the strongest answer.
3. Now try to DISPROVE that answer. Where is it weak? What would have to be true for it to fall apart?
4. Only if it survives, give me your recommendation, your confidence level, and the one thing most likely to break it.

If my leaning is flawed, tell me plainly. I want the truth, not reassurance.

Level 4 A loop

What it is

The eval check, but it keeps going. Work, grade, fix the weak part, repeat, until it is genuinely good or hits a stop you set.

Reach for it when

The first draft is never good enough and you want quality, not just done.

Cost higher

A loop comes in three plain-English shapes. Start with the first and stay there whenever you can.

Checks itselfOne agent does the work and marks its own homework, round after round. The common default, and all most jobs need.
Maker plus checkerOne agent makes it, a second separate agent grades it. Reach for it when the judging is a matter of taste, not a countable rule.
Manager with helpersOne agent runs the show and hands parts to helper agents under it. The most power, the most ways for it to break. Rarely needed.
Copy, paste, fill the brackets
Do this in cycles until it is genuinely good, not just finished.

The task: [WHAT TO PRODUCE]
What "good" means here: [THE STANDARD, BE SPECIFIC]

Each cycle:
1. Produce or improve the work.
2. Grade it honestly against the standard above. List the weak spots.
3. Fix the weak spots.
4. Stop when it clearly meets the standard, or after [NUMBER] rounds.

Show me the final version, then a short note on what you changed and why.

The full method, the three shapes, and worked examples are in the loop.

Level 5 /goal

What it is

A loop with a finish line built in. You name a target it can pass or fail, and it keeps going until it passes.

Reach for it when

You can name a clear finish line and want it chased without hand-holding.

Cost high
Copy, paste, fill the brackets
Work towards a goal and grade yourself until you hit it.

Goal: [THE OUTCOME, STATED SO IT CAN BE PASSED OR FAILED]
Pass test: [HOW WE KNOW IT IS DONE, the bar to clear]
Constraints: [ANYTHING FIXED: time, scope, do-nots]

Then:
1. State your plan to reach the goal.
2. Do the work.
3. Grade the result against the pass test. Give it a clear pass or fail.
4. If it fails, fix what failed and grade again. Repeat.
5. Stop when it passes, and show me the final result plus the grade.
Put it on a clock: scheduled tasks · runs without you
  • Describe a recurring job in plain words and Cowork runs it on a schedule: no reminding, no being there.
  • One catch worth knowing: it only runs while your computer is awake and the Claude app is open. A missed run goes late, not never, so do not lean on it for anything time-critical.
  • Schedule a task only after you have proven it works by hand. Automating an unreliable task just multiplies the mess.
At work: an operations lead schedules: "every Monday at 8am, pull last week's support emails, group them by issue, and give me a one-page summary of what needs escalating." It is waiting on their screen when they sit down.
Set a scheduled task
Set this up as a scheduled task that runs on its own.

WHEN: [e.g. every Monday at 8am]
WHAT TO DO EACH RUN: [the job, step by step, in plain English]
WHAT I WANT BACK: [the exact output, e.g. a one-page summary saved to my drive]

Draft the scheduled task and show it to me to approve before you turn it on. Note: I understand it only runs when my computer is awake and the Claude app is open.

Level 6 Ultracode

What it is

A whole team of agents working at once, splitting the job and checking each other. The manager-with-helpers loop, scaled up to build real software end to end.

Reach for it when

The job is too big for one worker, or you want real software built from start to finish.

Mission Control: a live dashboard showing an orchestrator fanning a mission out to a team of agents, with checks landing in lanes below.
You cannot usually see a team of agents work. This is what it looks like: one mission fanned out, checks landing live. Mission Control →
Cost highest
Copy, paste, fill the brackets
Run this as a team, not a single worker.

The job: [THE BIG OUTCOME]
Done means: [WHAT THE FINISHED THING LOOKS LIKE]

Plan it as a team:
1. Break the job into parts that can run in parallel.
2. Assign each part to its own agent with a clear brief.
3. Have one agent review and join the parts into a single whole.
4. Check the joined result against "done" above, and fix any gaps.

Show me the plan and who does what before you start.

This is the biggest of the three loop shapes. The deep dive on all three is in the loop.

Inside the team: sub-agents · how the splitting works
  • Each sub-agent works in its own clean, separate context, so none get bogged down, then they report back and the results get combined.
  • They work alone and cannot talk to each other, so this only fits work that splits into genuinely independent chunks.
  • Anthropic's own test found a team of sub-agents beat a single agent by 90.2% on a research task, at roughly fifteen times the cost. Power and bill climb together. See the study.
At work: a manager prepping a quarterly review: one sub-agent pulls the numbers, one summarises customer feedback, one checks what competitors shipped. They run at the same time, then hand back one combined brief.
Beyond the six · the frontier, rarely needed
  • The step past this page: the AI writes a small plan, then fans out a large parallel job, sometimes dozens or hundreds of helpers at once, and merges everything at the end.
  • /goal runs one job deep until it is right. This runs many jobs wide at the same time.
  • It always asks before running. You will not trigger it by accident.
  • Honestly, you will rarely need it. One person burned through half a $200 monthly plan with a single prompt at this level. If a normal chat, a skill, or /goal solves it, use that instead.
At work: a founder points it at a shared drive of hundreds of contracts and runs one pass that reads and tags every file at once, then hands back a sorted, searchable index. Genuinely useful, and clearly overkill for a normal day.
Check the cost before you go this big
This may need a big parallel workflow. Before running anything, tell me:
1. Does this job actually break into many independent pieces that can run at the same time? If not, recommend a cheaper approach (plain chat, a skill, or /goal) instead.
2. Roughly how much will this cost in effort/tokens, and is it worth it?

If a big parallel run really is the right call:
THE JOB: [the big parallel task]
THE DELIVERABLE: [one clear, named output]
Put the helpers on the cheapest capable model, bound the scope tightly, and ask me to confirm before you run it.
Four habits that make every level work · with the evidence

1 · Stop the yes-man. AI fails to push back about 88% of the time, versus about 60% for people, so make disagreement its job. The receipt.

Make it push back
Before you agree with me, argue against me.

MY POSITION: [what you think / want to do]

1. Give me the single strongest case AGAINST this.
2. Tell me what I'd have to be mistaken about for my position to fall apart.
3. Then, and only then, give me your honest view, even if it contradicts me.

Don't soften it to be nice. I'd rather hear the flaw now than have a client or my boss find it later.

2 · Finished is not working. "Done" is a claim, not a proof, so ask for the evidence it is done, not the announcement. The receipt.

Prove it's done
You've told me this is done. Now audit it as a hostile reviewer who WANTS to prove it isn't finished and gets nothing for being kind.

1. List the things that had to be true for this to count as finished.
2. Go through each one as that hostile reviewer and show me, concretely, that it holds. Where you can't show it, mark it FAILED, not "probably fine."
3. Tell me the one part most likely to be off, thin, or skipped.

If any box isn't genuinely ticked, say it's not done yet rather than papering over it.

3 · Mind the dumb zone. A long session gets sloppy long before it runs out of room, so hand off to a fresh chat with a tight summary. The receipt.

Clean handoff to a fresh chat
This chat is getting long. Write me a tight handoff I can paste into a brand-new chat so we lose nothing important:
- What we're doing and why
- Decisions already made (and what's settled, so we don't reopen them)
- The current state / where we got to
- The exact next step

Keep it short and concrete. No recap of the chit-chat.

4 · Pick the lowest level. That one is the golden rule at the top of this page, and its prompt sits right under it.

Do not want to memorise the ladder?

A skill can pick the level for you. Describe your task in plain words and it reaches for the lowest level that does the job. Meet the translator skill.

Check the words stuck

You have seen every core term on this page. Flip through them on the Vocab gym →