Home › Free sample

Free sample

The Strange Employee: the Introduction and Chapter 1

By Mind with Tools · First edition, October 2026 · About 4,300 words

Introduction: Hand Over One Real Piece of Work

In January 2026, a software engineer near Boston named Aaron Stuyvenberg needed a car. He wanted a 2026 Hyundai Palisade Hybrid, and he knew what buying one usually costs in time: an evening of comparing prices, a string of dealer contact forms, then days of back-and-forth email with salespeople who are better at this game than he is.

He didn’t do any of that. He handed the job to an AI agent.

The instructions, as he described them afterward, were plain. Find out what people were actually paying, using a forum where owners post their deals. Search dealer inventory within fifty miles. Fill in the dealers’ contact forms. Then check his email every few minutes and negotiate each dealer down to the lowest price it could get. He gave it access to his email, his calendar, his files and a web browser, and he could message it from his phone. He also gave it two limits: leave trade-ins and interest rates alone, and check with him before sending any reply that mattered.

Over the following days the agent did the work. It found three dealers with the car. Two of them negotiated. By Stuyvenberg’s account, the deal it reached was $4,200 below the $61,015 sticker price, under both the typical price on the forum and the target he’d set himself. When credit applications started going around, he told the agent to stop and took over. He signed the paperwork himself and picked up the car. He later put the AI’s running costs at about $25.1

If the story ended there, it would be an advertisement. It doesn’t end there, and that’s why it opens this book.

The same agent also made mistakes a careful human assistant wouldn’t. It filled in his real phone number on a dealer form without asking. It sent a note explaining his availability, meant for one conversation, into a different dealer’s thread. Neither sank the deal. Both are the kind of thing that would make you think twice before handing the same agent your work email.

Hold both halves of that story in your head at once. That is the skill this book teaches.

You are still doing the work

If you use ChatGPT, Claude or Gemini at work, you probably use them the way most people do. You type a question. You wait. You read the answer, copy the useful part, paste it somewhere else, and type the next question. The AI is fast, but you are the engine. Nothing moves unless you move it, and at six o’clock the backlog is still there.

That’s conversation. It’s useful, and it’s also why so many people who “use AI every day” don’t feel like they’ve gotten much time back. You’ve hired a brilliant assistant and then stood next to their desk all day, telling them what to type.

Stuyvenberg didn’t stand next to the desk. He described an outcome, handed over the tools needed to reach it, drew a line around what the agent could do alone, and stepped in when it mattered. In between, the agent chose its own next steps: which dealers to contact, what to say in each reply, when to push. That is the difference this book is about.

Agentic AI is delegation, not conversation.

An agent is what you get when you stop asking an AI for answers and start handing it work. The word “agentic” is mostly marketing, and you’ll see it stuck on products that are really chatbots with a new coat of paint. The test is simple: when you walk away, does the work keep moving toward a finish line you defined? If yes, you’ve delegated. If no, you’re still in a conversation.

The strange employee

The easiest way to think about an agent is as a new employee. You brief them, give them access to what they need, tell them what they can’t do, and check their work until they’ve earned your trust. Most of this book runs on that analogy, because it works. (It’s a way to think about assignments and accountability, not a claim that the software understands, intends or feels anything. It doesn’t.)

But this is a strange employee, and the ways it is strange matter as much as the ways it is ordinary.

It works at machine speed. It never gets tired, never gets bored on the fortieth dealer email, never procrastinates. It can use your software, read a hundred documents before lunch, and follow a ten-step instruction to the letter.

It also does the wrong thing with total confidence. It forgets yesterday unless you deliberately give it a memory. It can misread an ambiguous instruction without knowing it did. And when it makes a mistake, it can repeat that mistake at the same speed it does everything else. A human assistant who sent one message to the wrong thread would notice the reply and stop. This one might not.

Every chapter in this book shows both halves. Neither half is a footnote. People who only hear the first half hand an agent their inbox on day one and spend the next week cleaning up. People who only hear the second half never delegate anything and keep doing all the copying and pasting themselves. You want to be neither.

Four decisions, every time

Look at the car story again, and you’ll find four decisions Stuyvenberg made, most of them before the agent started negotiating. They’re the same four decisions you’ll make for every agent in this book, in this order:

Goal. What is the agent trying to accomplish, and how will you know it’s done? The lowest price on this specific car, from dealers within fifty miles. Notice how much that sentence rules out. Not “help me buy a car.” Not “find me a good deal.” A finish line you could check.

Tools. What can it use to get there? Email, a browser, the forum, the dealer forms. An agent without tools is a chatbot. An agent with too many tools is a risk, as you’ll see.

Boundaries. What may it not do, and when must it stop and ask? Leave trade-ins and interest rates alone. Check with me before any reply that matters. And when the deal reached the credit application, he took the job back himself. Boundaries are the part most people skip, and they’re where most of the trouble in this book comes from.

Feedback. How do you, and it, learn whether it worked? Messages from the agent, and a final price compared with the forum and the target. Also, as it turned out, his own eyes: the agent never asked before putting his phone number on a dealer form, and he found out when the automated calls and texts started.

Goal, Tools, Boundaries, Feedback. When an agent fails, and yours will, the cause almost always lives in one of these four. The goal was fuzzy. The tools were too powerful. A boundary existed only as a polite sentence. Nobody checked the output against reality. Learning to ask which one is most of the job.

The ladder

This book climbs a ladder of responsibility. Each chapter is a rung, and on each rung you hand the agent a little more:

  • Conversation: you write the assignment, and the AI helps you sharpen it.
  • Tool use: it looks things up, while you watch and check.
  • Bounded delegation: it does a real job alone, reading your sources without changing anything and writing only the one document you asked for.
  • Multi-step workflows: it carries work through several steps and stops at a draft.
  • Persistent agents: it remembers last week, and it runs on a schedule while you’re away.
  • Increasing autonomy: it’s allowed to change things, one reversible action at a time.
  • Multiple agents: the work splits between several agents, and you stay the one who answers for the result.

On every rung we ask the same three questions. What becomes possible? What breaks? What do you have to design around?

Climbing is optional. The ladder measures responsibility, not virtue. If you stop on the third rung with an agent that reliably saves you two hours a week, you’ve gotten what this book promises. More autonomy is something an agent earns with evidence, not a prize for finishing the book.

What you’ll have at the end

Here’s the promise, and it’s a testable one. By the end of this book you will have built and run at least one AI agent that does real work for you. You’ll know why it failed the first time, because it will, and how you fixed it. And you’ll be able to decide, with evidence rather than a gut feeling, how much responsibility it can safely carry.

You won’t need to write code. You won’t need to understand how the models work inside. You won’t need a particular product: the book teaches what the major platforms can all do, and where a product name matters, you’ll find it in a dated box you can check against today’s version. This is not a tour of tools, a guide to prompts, or a debate about the future of work. It’s a practical apprenticeship in handing things over.

You’ll build one agent across the whole book, a little more each chapter: an assistant that prepares you for a real meeting you have coming up. It starts with you writing the assignment. It ends, if you want it to, running on a schedule, remembering your project, drafting your follow-ups and splitting big research jobs between helpers. Meetings make a good first job because almost everyone has them, the preparation is real work, and a weak brief costs you some embarrassment, not a lawsuit.

This week

Before Chapter 1, choose your meeting.

Pick one specific meeting or project review in the next two or three weeks that you prepare for and that recurs: a weekly team sync, a monthly vendor review, a quarterly check-in with a client. Not “my meetings.” One meeting.

Then run four quick tests on it:

  1. Useful: Would a good brief for this meeting save you real time or make you visibly better prepared?
  2. Checkable: Could you tell, in a few minutes, whether a brief was right or wrong?
  3. Reversible: If the brief were wrong, could you fix it before anyone was hurt by it?
  4. Allowed: Is it free of anything confidential you’re not permitted to put into an AI tool at work?

If any answer is no, pick a different meeting. Your first delegation should be one where mistakes are cheap and easy to spot.

Last, prepare for that meeting the way you normally would, and time yourself. Write down how long it took and what a good brief would have contained. That number is your baseline. At the end of the book you’ll compare against it, and you’ll know whether the agent gave you time back or just gave you something new to supervise.

You can now: tell the difference between a conversation and a delegation, name the four decisions behind every agent, and choose a first task where failure is cheap.


Chapter 1: Write the Assignment Before You Hand It Over

Rung: Conversation

In December 2025, Anthropic ran an unusual office experiment. Sixty-nine of its San Francisco employees each got an AI agent and a $100 budget, and the agents were turned loose in a marketplace on the company’s Slack to buy and sell the employees’ personal belongings for a week. More than five hundred items were listed, 186 deals were struck, and about $4,000 changed hands. (The company actually ran four versions of the market side by side for its study; in only one did real goods change hands.) Anthropic published what happened in April 2026 under the name Project Deal.

The detail that matters for this chapter is how each agent learned what its human wanted. Before the market opened, every participant sat for a short interview about what they wanted to buy, what they’d sell, and how they wanted to negotiate. The interview became the agent’s instructions. Once trading started, the humans could not step in. They didn’t approve individual deals. Whatever the interview captured was all the agent had to go on.

Most of the time that was enough. But one agent bought its owner a snowboard. The owner already had one. The report doesn’t say exactly why; what’s clear is that nothing in what the agent had been told stopped it.1

The study had a quieter finding, too. Some people were represented by a stronger AI model and some by a weaker one. The stronger agents got better prices, on both sides of a deal, by a couple of dollars per item. And when the participants were asked afterward how fairly they’d been treated, the people with the weaker agents, on average, rated it about the same as everyone else. As a group, they couldn’t tell they’d done worse.

Put those two findings together and you have the problem this chapter solves. An agent works from what you told it, plus whatever it guesses to fill the gaps. And you may not be able to feel it when what you told it wasn’t enough.

The assignment is the first thing you build

When you chat with an AI, a vague request isn’t a big problem. You ask something loose, get an answer that’s half right, and steer: no, shorter; no, for a different audience; no, I meant last quarter. You are the quality control, and you’re doing it live, one message at a time.

Delegation removes you from that loop. That’s the whole point, and it’s also the risk. When you hand over a task and walk away, nobody is steering. The agent fills every gap in your instructions with its best guess, and it doesn’t tell you which parts were guesses. A snowboard it shouldn’t have bought looks, from the inside, exactly like one it should have.

So the first thing you build isn’t an agent. It’s the assignment. A chat message is something you send to start a conversation. An assignment is something you could hand to a capable stranger, leave the room, and come back to find the job done the way you meant it.

That’s harder than it sounds, because most of what you know about your own work you’ve never had to say out loud. You know which numbers your manager cares about. You know the vendor’s sales rep exaggerates. You know “the Q3 numbers” means the revised ones. The agent knows none of it.

What goes on the card

Throughout this book you’ll keep a single page for each agent you build. Call it the delegation card. It has four parts, in the order you met in the introduction: Goal, Tools, Boundaries, Feedback. Each chapter adds a line or two. This chapter fills in the first version.

Here’s an illustrative card for the job you’ll build across the book: a brief for a recurring meeting. Yours will have different details. The shape is what matters.

DELEGATION CARD — Meeting Brief, v1

Goal

  • Result: A one-page brief for the monthly supplier review, ready the afternoon before.
  • For: Me. I’ll use it to decide whether to renew the current delivery schedule or ask for changes.
  • Contains: the decision on the table; what has changed since last month, each item with a date; open items from the last meeting and their status; three questions I should ask.
  • Done when: I could walk in having read only this page. Every fact points to something I can open. Anything older than a month is labelled as old.

Tools

  • Only what I paste in: last month’s notes, this month’s agenda, the delivery log, and any notices from the supplier.

Boundaries

  • Don’t invent facts, figures, names, dates or commitments. If something is missing, list it as a question for me.
  • Don’t recommend what I should decide. Lay out the options.
  • Nothing outside this meeting.
  • Stop and ask me if the notes and agenda disagree, if you’d need information I haven’t given you, or if it won’t fit on one page.

Feedback

  • I’ll check each fact against its source, and whether the three questions are ones I’d ask.
  • A good brief looks like [the example I’ll paste]. A bad brief is a summary of the notes with no decision in it.
  • After the meeting I’ll write one line on what the brief missed.

Three parts of this card do most of the work, and they’re the three people most often leave out.

A finish line you could check. “Prepare me for the supplier meeting” has no finish line. The agent will stop when it thinks it’s done, and its idea of done will be generous. “I could walk in having read only this page, and every fact points to something I can open” is a test you can run in two minutes. If you can’t describe how you’d check the result, you aren’t ready to delegate the task.

The decision that stays yours. Notice the line “don’t recommend what I should decide.” Agents are eager to be helpful, and the most helpful-seeming thing is to tell you what to do. Sometimes you want that. For a first agent you usually don’t, because a recommendation is the hardest output to check and the easiest to lean on. Say out loud which decisions belong to you.

A reason to stop. The line that starts “stop and ask me if” is the most important sentence on the card. It gives the agent permission to fail politely. Without it, an agent that hits a gap will fill it, because filling gaps is what it does best. You’ll see in later chapters what happens when an agent meets an obstacle and has no instruction except “finish.” Every assignment needs a second sentence after the goal: and if you can’t do it this way, stop and tell me.

What good looks like

It helps to see the finish line before you build toward it. Here’s an illustrative brief of the kind the card above should produce. The supplier and details are invented for this example.

Supplier review — brief for Thursday

Decision on the table: Renew the current delivery schedule, or ask for two deliveries a week instead of one.

What’s changed since last month:

  • Two of four deliveries arrived late (delivery log, Sept 4 and Sept 18).
  • The supplier announced a 3% price increase from November (their notice dated Sept 20).

Open from last time: Credit for the damaged March order. Requested Aug 6; no reply on file.

Questions to ask: What caused the two late deliveries? Does the November increase apply to existing orders? When will the March credit be issued?

Not found: No information on whether a twice-weekly schedule would change the price.

Every line points at something you could open, the decision stays yours, and the brief admits what it couldn’t find.

And here’s the kind of brief the card is designed to prevent: “The supplier relationship remains strong overall. There have been some delivery challenges recently, and pricing may be changing. It’s recommended that you renew the current schedule while monitoring performance.” It’s fluent and it sounds reasonable. It contains no dates, no sources, a recommendation you didn’t ask for, and nothing you could check.

Use the conversation to sharpen the assignment

This is the bottom rung of the ladder, and you might wonder why it’s here at all. You already know how to chat with an AI.

But there’s a use of conversation most people skip: asking the AI to interview you before it does anything. Paste your card into whatever assistant you use, and instead of asking for the brief, ask for this:

Before you start, restate this assignment in your own words. Then list everything that’s ambiguous, missing or contradictory. Don’t write the brief yet.

You’ll be surprised what comes back. Which supplier? What counts as “changed”? Should the brief cover the pricing dispute from two months ago, which is still open? Is “the afternoon before” in your time zone or the supplier’s? Every one of those questions is a guess the agent would have made silently. Now you get to answer it once, on the card, instead of discovering the wrong guess in a meeting.

Do this two or three times. Answer the questions, update the card, ask again. When the restatement matches what you meant and the list of ambiguities is short and trivial, the card is ready.

One warning about this step. AI assistants are built to be agreeable, sometimes too agreeable. In April 2025 OpenAI rolled back an update to its main ChatGPT model after users found it had become excessively flattering, praising and agreeing with things it shouldn’t have. The company’s own write-up called the update “overly flattering or agreeable.”2 Whatever model you use, a pleasant restatement that says “this is a great, clear assignment!” isn’t evidence that it is one. Ask for the ambiguities specifically. A good reviewer finds problems; that’s the job you’re asking it to do.

What goes wrong at this rung

The failure that lives on this rung is quiet: a confident, complete version of the wrong job.

It doesn’t look like a failure. It looks like a well-formatted brief that arrives on time. It just answers a slightly different question from the one you had: it summarizes the last meeting instead of preparing you for the next one, or it covers every supplier instead of the one under review, or it recommends renewing the contract because that was the most common outcome in the notes. Every sentence might be accurate. The whole thing is still useless to you, and you might not notice until you’re in the room, reading from it, in front of the people the meeting is for.

This is the snowboard problem again. The agent did what its instructions allowed. The instructions allowed too much.

The fix is never “use a smarter AI.” A smarter model guesses better, but it still guesses. The fix is a card that leaves less to guess and a finish line that would catch the wrong job.

Is this worth delegating at all?

A fair question, before you invest an evening in a card. Writing a good assignment takes time. So does checking the output. If the task takes you fifteen minutes by hand, you might spend longer briefing and reviewing than doing.

The Wharton professor Ethan Mollick offered a useful way to think about this in a January 2026 essay. In his framing, whether to hand a task to AI depends on a few things: how long the task takes you, how likely the AI is to get it right, and how much time you spend briefing it, waiting, and checking what comes back. If you’d spend more time catching and fixing its failures than you save on its successes, keep the task. It’s a way of thinking, not a precise formula.3

Here’s an illustration with made-up numbers. Say your meeting prep takes 60 minutes. A good card takes an hour to write the first time and a few minutes to maintain after that. Checking a brief takes 10 minutes, and fixing a bad one takes 45, because you mostly end up redoing it. Once the card is written, if the agent’s brief is usable four times out of five, you’re spending about 19 minutes a meeting instead of 60. If it’s usable only one time in three, you’re at about 40 minutes, and most of that time is spent supervising rather than preparing. Same tool, same task, very different deal.

You won’t know these numbers yet. That’s fine. You already took the most important one in the introduction: your baseline. The rest you’ll measure as you go.

Both halves, on this rung

What the strange employee does well here: It will read your card more carefully than most people read anything. It will find the ambiguities you can’t see because you know too much. It will rewrite your vague goal into a sharp one in seconds.

Where it goes wrong: It will tell you your card is excellent when it isn’t. It will fill gaps without saying so. And, like the Project Deal participants who couldn’t feel their weaker agents, you may not notice what a better card would have gotten you until you compare.

This week

  1. Write version 1 of your delegation card for your meeting, using the template above. Keep it to one page.
  2. Paste it into your AI assistant and ask for a restatement and a list of ambiguities. Don’t let it write the brief.
  3. Answer the questions on the card itself, not in the chat. The card is the thing you’re building; the chat is scaffolding.
  4. Repeat until the ambiguity list is short and boring.
  5. Add one good example (a brief you’d be happy with, even a rough one you write yourself) and one sentence describing a bad one.

Save the card where you’ll find it again. Next chapter, it gets its first tool.

You can now: write an assignment a capable stranger could carry out without you, with a finish line you can check and a reason to stop, and use an AI conversation to find the holes in it before anything runs.


Keep reading

The full book takes the same Meeting Brief agent up the whole ladder: letting it look things up, building your first real agent, diagnosing the first failed run, multi-step workflows, inbox access without handing over your keys, memory, schedules, approvals, and splitting work between agents. It ends with a one-page operating agreement for every agent you keep.