"AI agent" has become a fuzzy term. In practice an agent is nothing more than a language model with access to some tools, running in a loop until the job is done.

A worked example

Say dozens of support emails arrive every day. We want an agent that will:

  1. Read the incoming email
  2. Classify it (technical / billing / sales)
  3. If technical, search the knowledge base
  4. Draft a reply
  5. Send the draft for human approval — not send it itself

Step five is the most important part, and the one most often skipped.

The shape in n8n

Email Trigger (IMAP)
      ↓
Extract & Clean  ── strip signatures, quotes, repeated text
      ↓
AI Agent Node
      ├── Tool: searchKnowledgeBase(query)
      ├── Tool: getCustomerHistory(email)
      └── Tool: checkOrderStatus(orderId)
      ↓
Switch (confidence)
      ├── high → draft into approval queue
      └── low  → straight to a human
      ↓
Slack / Draft mailbox

Practical notes

Keep tools small and specific

A good tool does one thing and returns something predictable. Instead of one doEverything(), write three separate tools. Models make better decisions with unambiguous tools.

Always ask for structured output

Rather than free text, have the model return JSON:

{
  "category": "technical",
  "confidence": 0.86,
  "draft": "…",
  "sourcesUsed": ["kb-142", "kb-097"]
}

Take sourcesUsed seriously. If the agent can't tell you where its answer came from, that answer isn't trustworthy.

Define a confidence threshold

Every agent is sometimes wrong. What matters is that it knows when it is unsure. Below whatever threshold you set, the work should go straight to a person.

Track cost from day one

Every run costs tokens. n8n can record consumption. Without it, three months later you meet a bill nobody expected.

A rule we hold to at Argbaan: no agent reaches production without full logging of its inputs, outputs and the tools it called.

The maturity path

Don't turn an agent loose in one go. Move through three stages:

  1. Observer — it runs and produces output, but takes no action. Compare its output against the human decision.
  2. Suggester — it drafts, a human approves or corrects.
  3. Bounded actor — it acts on its own for low-risk cases; high-risk ones still need approval.

Stage two usually delivers the most value for the least risk. Plenty of teams never need stage three at all.