field-report

Field report. First written in July 2026, rewritten in October 2026.

Why I run my company with AI agents

In my own words: how I got here, how I work with the agents today and the one thing they're not allowed to do. At the end, the details of the system underneath.

How it started

I've used AI since it came out: checking a text, small jobs, a bit of help with an email. But around the end of 2025 the models started to get really smart, and suddenly the whole topic got exciting.

I started handing real work to AI agents for two reasons. They can process data on a completely different scale than a person can, and if you set them up right, they're much better at judging the big picture. And honestly, coordinating people across time zones, with constant turnover, was the hardest part of the business.

Whenever someone left, we had to train the next person from scratch. We worked with SOPs, but we're set up internationally and I don't always live in the same place, so time differences kept getting in the way. A lot of it ran on freelancers, and my team was spread around the world.

I started the way I trained my first employee

With assistant work: collecting receipts, sorting documents. Nothing with a lot of responsibility, nothing where much can go wrong. The AI was good at it and worked through the processes well. I caught the odd mistake, built safeguards for it, and they worked.

Some of those early safeguards I hardly need anymore: the models got that good in one year, and the big labs like Anthropic or OpenAI build a lot of it into the models themselves. The ones that matter I kept. You'll find them further down. I still keep adjusting things so the AI doesn't do anything stupid. It's been a journey. A lot has changed in the last year.

How I work today

Today I mostly work in the loop. I orchestrate the sessions, give the answers they need from me and make the decisions. The AI works through its own tasks, 24/7. My team used to sleep while I was working, because they were on the other side of the world, or I was. For me that's a big blessing.

I don't fill in forms or file applications anymore. The AI does that. I check the form, then I tell the AI it may put my signature on it. At that point it stops, and a window pops up on my Mac: do you allow this signature? I click yes, the AI signs and prepares the email. I read it, check the attachment once more (is it the right one, is everything in order?), and then it goes out. One command.

That's so cool, because the AI knows everything. No employee of mine ever had all the information. I do: it's stored with me, backed up and encrypted, and the AI has access to it.

A product launch used to be a huge job with different teams: someone to coordinate the native-speaker translations, someone for the keyword research, people for the images. Now I tell Claude: here's the new product. It already knows the product, because we set it up together, and it knows our image style, so it hands the images to Codex to create. Before, I needed several people for that. Today my e-commerce business hardly takes my time, because everything we've learned over the years is written down in files and databases the agents can read. It runs my Amazon advertising and helps me with sourcing. It takes the game to a whole other level.

The one thing it's not allowed to do

I still send every email myself. The AI writes it and polishes it, but no email goes out on its own. The AI is allowed to do quite a lot for me, but not that. I want to decide what goes out, because at the end of the day my name is on it.

The mistake that annoyed me most: once a wrong tax number went out in two emails, because I didn't check it. I thought, yeah, that'll be fine. Since then the AI is only allowed to copy company numbers word for word from one master file, and every draft is read by a second model in its own context before I see it.

Beyond the business

A specialist will rarely be able to explain to a layman what's going on in their field. My specialist agents can, in normal words. When I want to understand a topic, I have Claude put a whole team of them on it: they research worldwide, auditors check their work, and the agents debate it with a referee. A real consortium. In the end I get a clean research report that I actually understand. Every model has its blind spots, so I use models from different labs and countries: from the US, from China and local ones on my own Mac. That way the big picture is always there.

What I'd tell other entrepreneurs

Just start. Look at your own topics together with the AI. You have to interact with it and work through the processes together: see it as an assistant, a partner and a work machine. Depending on the business, it can already take over most of the work. People underestimate that.

Where this is going

Honestly, not much has surprised me, except maybe how good these models got in a single year. And I can see where this is going once it's connected to robots: a lot is still going to change.

It's an exciting time. I want to be among the people who understand it and are out in front, and maybe invent something that helps others, in research too. Instead of fighting it, I enjoy working on it. For me it's like a computer game.

The system underneath

For everyone who wants the details, this is how it looks in October 2026.

  • 119 background jobs. Almost all of them run without AI, on purpose.
  • 85+ of my own skills for Claude Code.
  • 20 subagents: advisors, reviewers and a researcher.
  • 10,000+ automated tests in the main test suite.

What the agents do for me

  • Email: when I ask, agents go through my five mailboxes and draft the replies. Every draft is checked by a second model and by fixed rules. Then it lands in a small window on my Mac. I read it and send it myself, with one click.
  • Advertising: we do our Amazon ads ourselves. Agents prepare the changes every week, and a second agent reviews them. Nothing that moves money happens without my OK.
  • Guests: for the holiday apartment I run in the Swiss Alps, the door code and a fixed message before arrival go out automatically through the booking system. My personal replies to guests are drafted by agents, and I send them.
  • Specialists: advisor agents for inventory, pricing, finance, account health, reviews, product research and advertising. Each one answers from its own knowledge base.
  • Everything else: invoices, FBA shipments to Amazon, listing checks and review requests.
  • Day and night: a watchdog looks at all background jobs. If one hangs, a simple repair step restarts it, without any AI. If that doesn't help, an investigator agent takes over.

How I keep it under control

  • New things only make suggestions at first. I approve, and only a small set of actions ever runs on its own. The investigator ran one week under my supervision first. In that week it correctly sorted out 17 false alarms and fixed 6 real problems after my OK. Nothing had to be undone.
  • Models from other labs check the work. Claude's code gets reviewed by OpenAI's Codex, Zhipu's GLM and Moonshot's Kimi. Claude checks every finding against the code before anything gets fixed.
  • Only emergencies reach my phone. Small problems with a safe fix get fixed automatically. The rest goes to a dashboard and into one email a day.
  • Every loop has an off switch. Before it changes something, it saves a copy. If the check fails, it goes back to that copy. After two tries without success, it stops and asks me. And no agent can change the rules it's checked against.
  • When something goes wrong, I make a rule out of it. A model once wrote "Friday" next to a date that was a Saturday. Since then a script checks the weekday of every date in my email drafts.
  • The AI has to call me by name, in every reply. When my name is missing, that's my sign the session has gotten too long and it's time for a fresh one.
  • I measure. A one-week check found a repair loop that changed nothing 95% of the time and still cost about $231 a week. The same day it got a simple check up front. In September I tested my email checks: one second reader plus fixed rules, against the four checkers I had before. The new way found more mistakes (92% instead of 82%), had the same rate of false alarms, and costs about 73% less. So I switched. It was a small test: 34 emails, most of them real, 17 with a mistake in them. The new way ran three times, the old one once.
  • AI only where it's needed. A script does the same thing every time. AI only comes in where something needs judgment: reviewing, sorting, writing drafts.

What went wrong

Of course things went wrong too. Here are four examples.

  • Live campaigns with paused ads. Three ad campaigns showed as live, but every ad group, keyword and product ad under them was still paused. I found it myself; none of my checks noticed it. Since then a script checks all four levels, from the campaign down to the product ads, before a campaign counts as live.
  • Four weeks of broken backups. A script that manages my backups did it the wrong way for four weeks before anyone noticed. Since then every script that touches files needs a test before it may run on a schedule.
  • Six fake alarms on my phone. A test sent its alarms to my real phone. Since then tests can't reach my real alerts.
  • Lost email drafts. A checker that was too strict threw away email drafts, and nobody saw it. We fixed it and added tests.

The stack

Backbone
Claude Code. A top Claude model for judgment and writing, smaller ones for routine steps.
Review
OpenAI Codex for code reviews, security scans and repairs. GLM (Zhipu) and Kimi (Moonshot) as extra code reviewers. Grok (xAI) to check plans.
Screening
TypeSafe Jev, a small AI model that only picks a category or gives a score. It checks incoming text for prompt injection and pre-filters ad search terms.
Research
Perplexity.
On my Mac
Qwen3, Whisper, BGE-M3 and SigLIP-2 for sorting, transcription, search and images.
Memory
Text files with everything we've learned, and Qdrant to search them.
Alerts
One email a day. Telegram only for emergencies.
Billing
Mostly flat-rate subscriptions instead of paying per token.

Open source

Lookout: a small Mac app that shows which AI session is working, which one is done and which one is waiting for you.

agent-ops-kit: safety rails for AI agents that run on their own.

Are you hiring for agent operations, or do you run your company with AI agents too? Let's talk on LinkedIn.