A row of work hats hanging on hooks along a warm off-white wall in soft daylight, with a telephone headset hung on a forest-green ribbon.

BUY BACK YOUR WEEK

A solo founder wears every hat in the building. Sales, bookkeeping, marketing, support, product, and the dozen small jobs that never make it onto a business card. Every AI vendor now promises help with all of them, and the promise is roughly half true. The expensive part is that nobody can tell you from the outside which half applies to your week.

Start with the finding that should unsettle anyone juggling nine jobs at once. In the largest field experiment yet run on AI at work, a team led by Fabrizio Dell'Acqua and Ethan Mollick studied 758 Boston Consulting Group consultants, most of them working with GPT-4 and a control group without. On most tasks the tool made the consultants who had it measurably better, and the effect was anything but subtle. Yet on a neighbouring task, chosen because the model would produce a wrong but convincing answer, the same consultants performed worse than colleagues working unaided, and almost none of them noticed which kind of task they were on. If elite specialists cannot feel the boundary between help and hazard, a founder wearing four hats before lunch has no chance of feeling it either. Something other than task difficulty is drawing that line, and finding it is worth a founder's attention.

📍 In brief

Hand AI the work you can mark faster than you could do it yourself. Keep the work you would have to redo in order to check. That one test sorts a founder's week better than any tool list, and the evidence for it is now solid enough to plan around.

The Wayfinder Team

Most of your load was never your job

The heaviest hats are the connective ones, and that is exactly where AI earns its keep first.

Ask a founder what their business does and they will name a craft, but ask where the week actually went and they will describe something else entirely. The hours go on chasing invoices, tidying the spreadsheet that feeds the cash forecast, turning meeting notes into a proposal, and writing the follow-up email that turns a good conversation into a booked order. None of it is the reason the business exists, yet all of it has to happen.

The usage evidence says this connective load is precisely what people hand to AI when they are free to choose. When Anthropic studied how knowledge workers were using its agentic Cowork product, roughly half of all usage turned out to be what the researchers called the work around the work, meaning tasks that appear in a broad swath of jobs but are rarely anyone's core responsibility. The two biggest categories, business operations and content production, are connective by nature. A spreadsheet pulls scattered numbers into one place where they can be compared. A deck carries a decision to people who were not in the room. A checklist moves knowledge from one head towards another.

For a solo founder the stakes on this finding are higher than for anyone in a big firm, because the person doing the connective work and the person whose judgement the business depends on are the same person. Every hour spent reformatting a quote is an hour taken from the only judgement the company has. So the first honest answer to the question in the title is unglamorous. AI lightens the hats you resent, not the hat that is your craft, and for most founders the resented hats are most of the week.

You cannot see the frontier from where you stand

AI ability is jagged, and its failures arrive looking exactly like its successes.

The consulting experiment gave the pattern a name that has stuck, the jagged frontier. Inside the frontier, consultants using the model completed 12.2 per cent more tasks, finished them 25.1 per cent faster, and produced work judged 40 per cent higher in quality. Outside it, on the task built to sit beyond the model's competence, accuracy fell from 84 per cent unaided to somewhere between 60 and 70 per cent with the tool. The model did not fail loudly, it produced an authoritative, well-structured, wrong answer, and skilled people signed it off.

A related experiment on recruiters points at the mechanism. Given a highly capable screening tool, experienced recruiters grew careless, stopped exercising their own judgement, and made worse decisions than peers given a weaker tool or none at all. Dell'Acqua called it falling asleep at the wheel, and the better the assistant becomes, the less reason anyone feels to stay awake.

Translate that to a founder's desk and the danger becomes plain. The risk is not that AI fails on your bookkeeping or your product copy. The risk is that it fails fluently, on a Tuesday, while you are grateful for anything that looks finished. Tasks that look identical from the outside sit on opposite sides of the frontier, and the research is blunt about the fact that difficulty does not predict which side is which. So what does?

Checking, not doing, is the real cost line

A task belongs with AI when you can mark its output faster than you could produce it, and it stays with you when checking means redoing.

The analyst Benedict Evans has the clearest account of where generative AI found its footing first, in the drafted email, the product copy, and the code you can run and test. In each of those a mistake is cheap to see and cheap to fix, so the error rate barely matters, and the time saved is real. Then he names the opposite class, the tasks whose answer is binary, right or not right, in a domain where you are not the expert. There, verifying the output means repeating the entire job yourself, and his conclusion is that you cannot usefully hand such work to a model at all.

Put those two classes side by side and the founder's sorting rule falls out. Call it the check-to-do gap. When the cost of checking a piece of work is small next to the cost of doing it, handing it over is nearly pure gain. When the two costs converge, the tool has added a layer of management to a job you still effectively own.

The rule is concrete enough to use on a real week. Marking a machine-drafted proposal against your own knowledge of the client takes minutes, because you would recognise a wrong claim on sight. Checking machine-prepared tax figures in a regime you do not understand is an accountant's job you would now be doing twice, badly. Same tool, same founder, opposite verdicts, and the deciding variable was never the software.

There is a genuine sweetener in the evidence for anyone sorting their hats this way. The consulting study, along with parallel work on writers, lawyers and support agents, found that AI helps weaker performers far more than strong ones. In the BCG data the bottom half of consultants gained 43 per cent in output quality against 17 per cent for the top half. A founder is by definition a bottom-half performer in most of their hats, because nobody is elite at nine jobs. The hats that fit worst stand to gain most, provided the check stays cheap.

A two-by-two grid comparing cost to do against cost to check. Dear to do and cheap to check: hand it over and mark everything. Dear to do and dear to check: stay in the loop or buy expertise. Cheap to do and cheap to check: automate it fully or stop doing it. Cheap to do and dear to check: keep it yourself.

The check-to-do gap: hand over what you can mark, keep what you would have to redo. Drawing on Dell'Acqua et al.'s field experiment and Benedict Evans's analysis.

A standing gate keeps help from becoming a layer

The founders who keep the lightness all share one habit, marking every machine-made piece of work before it leaves the business.

Anthropic's broader economic research shows where working habits are heading. In its study of embedded, programmatic use of Claude, 77 per cent of transcripts showed full delegation of the task, against 12 per cent showing genuine back-and-forth collaboration. Conversational use still splits evenly, but the direction of travel is towards handing work over whole. Delegation is where the productivity lives, and it is also where unwatched errors live, because delegation is precisely the mode in which nobody is awake at the wheel.

Our own operation is built on the assumption that both of those facts are permanent. Research, drafting and quality scoring run without a person in the loop, and every draft is scored against a bar before a human ever sees it. Yet nothing reaches a reader until a person has read it and pressed the button. That gate was not an afterthought or a regulatory nicety. It is the design decision the whole system leans on, and it consumed more build effort than the generation ever did, because producing plausible work is now cheap and trusting it is not.

The gate is also what makes the jagged frontier survivable for someone who cannot map it. You do not need to know in advance where the model's competence ends if every piece of work crossing the boundary passes your desk on the way out. The frontier stops being a hazard and becomes a line item, a few minutes of marking per delegated task, priced into the gain.

Where this doesn't apply

The sorting rule assumes you could mark the work, and sometimes you cannot. In domains with hard right answers that you would not recognise, tax, legal, anything regulated, cheap checking is an illusion, and the honest options are staying out or buying human expertise. The rule also stops at the work that is you. The sales conversation, the judgement call on a difficult client, the relationship that holds the business together. Handing those over lightens nothing, because they are not load, they are the product.

The bottom line

Do not start with the tools. Start with a list of this week's hats, and against each one ask a single question. Could you mark AI's version of this faster than you could do it yourself? Hand over the yeses under one standing rule, that you mark everything before it ships. Keep the noes without guilt. Buying a tool is the tactical move, installing the gate is the structural one, and the structural move is the one that compounds.

Give AI the work you could mark, and keep the work you would have to redo to check.

P.S. Watch the delegation numbers. As agent-style tools pull everyday use towards the 77 per cent pattern, the gap will widen between founders who kept a marking habit and founders who stopped checking. The second group will eventually locate the jagged frontier the expensive way.