How to write a phone agent script

Write spoken scripts for two-way calls — not chatbot turns. Persona, facts, escalations, and what not to promise.

Most phone agent scripts fail for the same reason: they're written like chatbot conversations, not like speech. Someone drafts tidy paragraphs where the caller says exactly the right thing and the agent replies with a polished monologue, and then it hits a real call — where the caller interrupts, mumbles a number, changes their mind, and asks two things at once — and the whole thing falls over. Writing for a two-way, realtime phone agent is a different craft. This guide teaches it, using Agent Vani as the reference, so what you write survives contact with real callers.

The core mindset shift: you're not writing a transcript. You're writing a persona plus a set of facts plus a set of boundaries, and trusting the agent to hold a natural conversation within them. Over-scripting every line makes the agent robotic; under-specifying the facts makes it vague or wrong. The skill is knowing what to pin down and what to leave to the conversation.

Write for speech, not for chat

Read every line out loud before it goes in. If it sounds like something a person would say on the phone, keep it. If it sounds like a written FAQ answer, rewrite it. Spoken language is shorter, plainer, and more forgiving than written language.

  • Short sentences. A caller can't re-read a long one. Break it up.
  • Plain words. "We're open till 8" beats "our operational hours conclude at 20:00."
  • No walls of text. If the agent would say five sentences where a human says one, cut four. Callers interrupt long replies anyway — and a good agent should let them.
  • Contractions and rhythm. "You'll," "we're," "that's." It's how people talk.

This matters more on the phone than anywhere else because the caller can't skim. Every extra clause is a second they're waiting instead of getting their answer. And because the agent is two-way — callers can cut in — a tight reply that can be interrupted beats a long one that steamrolls them. Why that back-and-forth is the whole point is covered in two-way realtime vs turn-based AI calls.

Start with the persona

Before any facts, define who's answering. The persona is the through-line that keeps every reply consistent, so you're not scripting each answer individually.

  • Tone. Warm and casual? Crisp and professional? Match your brand and your callers. A clinic front desk and a nightclub line should not sound the same.
  • Language. Agent Vani answers in Hindi, English, and Hinglish. Decide the default and let it follow the caller's lead. Don't force a caller into stilted English to be understood — that alone breaks the "live" feel. See multilingual phone support.
  • Identity and honesty. Give it a name if you like, but don't have it pretend to be a specific human or hide that it's an AI. Honesty is part of caller trust — see privacy and caller trust.
  • Boundaries of voice. How does it handle an annoyed caller? A confused one? Set the disposition once, in the persona, and it carries through the whole call.

A strong persona means you write far fewer literal lines, because the agent knows how to say things without being told each time.

Then load the facts it answers from

A phone agent is only as good as what you give it. This is the part to be exhaustive about, because vagueness here becomes vagueness on the call. Cover the things callers actually ask:

  • Hours — including holidays, lunch breaks, and "are you open today" edge cases.
  • Location and directions — landmarks, parking, which entrance.
  • Policies — returns, cancellations, deposits, what's included, what's extra.
  • Products and services — what you offer, in plain terms, with the details callers ask about.
  • Common questions — pull a week of real calls and write down what people actually ask. That list is your knowledge base.

Write facts as facts, not as scripted sentences. Give the agent "Return window: 30 days, unopened, receipt required" and let it phrase the answer conversationally. Don't write out "Thank you for asking about returns. Our return policy allows…" — that's the chatbot trap again.

The honest constraint: the agent answers from these facts and won't invent policy. That's a feature — it won't promise something you didn't approve — but it means a thin knowledge base produces thin, "let me have someone call you back" answers. Time spent here is the single highest-leverage thing you'll do.

Define what it must NOT promise

This is the section most people skip and most regret skipping. Explicitly list what the agent should never commit to:

  • Prices or discounts you haven't approved. "I can't quote that, but I'll have someone confirm" is the right answer, not an invented number.
  • Guarantees, timelines, or outcomes the business can't stand behind.
  • Anything regulated — medical, legal, or financial advice. A clinic agent handles scheduling and FAQs, not diagnoses. See clinics and healthcare front-desk.
  • Actions it can't actually take. Don't let it say "done, I've processed that" for something it can't do. Say what will happen instead.

A good rule: if getting it wrong would cost you money, trust, or a legal problem, it goes on the do-not-promise list and, where appropriate, becomes a handoff trigger.

Design the escalation moments

Handoff isn't a failure state — it's part of a good script. Decide, in writing, when the agent should stop and get a human, and script the transition so it feels smooth, not like hitting a wall.

  • Name the triggers. An upset caller, a request outside the knowledge base, an explicit "I want to speak to a person," a high-value or sensitive request.
  • Script the handoff line. Something honest and warm: "Let me get one of our team to help with that." Not a cold "transferring you now" click.
  • Pass the context. The agent should hand off what it already gathered so the caller doesn't repeat themselves — the thing IVR never did.

Full treatment in how to set business hours and escalation and human handoff and escalation.

Write the pieces that carry every call

A few moments recur on every call. Get these right and the rest follows:

The greeting. Short, branded, and it sets the language. "Thanks for calling [business], how can I help?" Not a monologue. The greeting is the first three seconds where the caller decides if this feels live.

The clarifier. When a caller is vague, the agent needs a natural way to narrow down — "Sure, is this about an existing order or a new one?" — without sounding like a menu.

The confirmation. For anything that matters — a spelled name, a number, a date — the agent should read it back. "So that's the 15th, at 4pm — correct?" This is where accuracy lives.

The graceful "I don't know." The agent will hit questions it can't answer. Script honesty, not bluffing: "I don't have that in front of me, but I'll make sure someone gets back to you." A confident wrong answer is far worse than an honest "I don't know."

The close. A clean ending that confirms what happens next. "You're all set — anything else?" then a warm sign-off.

Test it on a real call, not on paper

A script that reads well on paper can still fail on the line, because paper doesn't interrupt you. Before going live, dial the number and try to break it the way a real caller would:

  1. Interrupt it. Does it stop and listen, or finish its scripted line over you?
  2. Be vague, then get specific. Does the clarifier feel natural or menu-like?
  3. Ask something off-script. Does it fall back to honest "I don't know" and handoff, or invent?
  4. Change a detail mid-call. Does it track the correction?
  5. Ask for a human. Does the handoff feel warm and carry context?

Iterate on what you hear, not on what you imagined. This test loop is the heart of the going-live checklist.

Writing for interruption

Here's the thing that separates a script written for a phone agent from one written for a chatbot: on the phone, the caller will interrupt, and your script has to expect it. A two-way, realtime agent yields when the caller cuts in — which means long, uninterruptible monologues aren't just tedious, they're written for a medium that doesn't exist. Assume every reply might be cut short and design accordingly.

Practically, that means front-loading the important part of every answer. If a caller asks whether you're open, "Yes, till 8 tonight" comes first — the details about weekends can follow and can safely be interrupted. It means avoiding replies that only make sense if the caller hears all of them. And it means the persona should be comfortable being cut off: an agent that gets flustered or resets when interrupted feels broken, while one that gracefully stops, listens, and picks up the thread feels human. You can't fully script this, but you can write facts and a persona that hold up when the neat turn-taking of paper meets the messiness of a real call. The why behind this is the pillar guide, two-way realtime vs turn-based AI calls.

Handling the messy caller

Real callers are messy, and a good script anticipates the mess rather than assuming the tidy path. Some patterns to plan for:

  • The rambler. Some callers tell you their whole story before getting to the point. The persona should be patient, then gently focus: "Got it — so the main thing you need is the refund status, right?"
  • The two-in-one. "What are your hours and do you deliver?" The agent should catch both and answer both, not drop one. Writing facts (rather than one-question-one-answer scripts) is what makes this possible.
  • The corrector. "The 15th — no wait, the 16th." The agent should follow the correction naturally. Confirmation lines ("so that's the 16th?") catch these.
  • The unsure caller. Someone who isn't sure what they need. A warm clarifier that offers a couple of natural options beats a rigid "please state your reason for calling."
  • The frustrated caller. Sometimes the right script is a short one that gets them to a human fast. Not every call is one to win single-handed.

You don't write a separate script for each of these. You write a persona flexible enough to handle them and facts complete enough to answer them, and you test against them.

Localising the script for your callers

If your callers speak more than one language — and on many business lines they do — the script has to account for it beyond just "supports Hindi and English." Think about the natural mixing: many callers move between Hindi and English within a single sentence, and the agent should follow rather than forcing a choice. Think about the terms your callers actually use for your products and services, which may be colloquial rather than the formal names in your marketing. And think about tone across languages — warmth and formality don't map identically, so a persona that feels right in one may need a slightly different touch in another.

The goal is the same as everywhere else in this guide: the caller should be able to talk the way they naturally talk and be understood, without shifting into a stilted, computer-friendly register. That's a large part of what makes a call feel live rather than mechanical. More on this in multilingual phone support and global teams, local languages on the line.

Common mistakes to avoid

  • Over-scripting. Writing every literal line makes the agent robotic and brittle. Give persona plus facts plus boundaries, then let it converse.
  • Under-feeding facts. A thin knowledge base means "I'll have someone call you back" to everything. Do the boring work of writing down what callers ask.
  • Writing chat, not speech. Long, formal paragraphs that no one would ever say out loud.
  • No do-not-promise list. The fastest way to an expensive mistake.
  • Treating handoff as failure. It's a designed feature. Script it warmly.
  • Skipping the live test. The only test that reveals what real callers will hit.

Frequently asked questions

How long should a phone agent script be? There's no page count. Nail a strong persona, a thorough set of facts, and clear boundaries and handoff rules. Depth in the facts, brevity in the spoken lines.

Should I script every possible response? No — that's the chatbot trap and it makes the agent robotic. Define the persona and facts and let the agent hold a natural conversation within them. Script only the recurring moments: greeting, clarifier, confirmation, honest "I don't know," and close.

What if the agent doesn't know an answer? Script an honest "I don't have that, but I'll make sure someone gets back to you" and, where it matters, a handoff. It answers from approved facts and won't invent policy — an honest "I don't know" beats a confident wrong answer.

How do I stop it from over-promising? Write an explicit do-not-promise list: unapproved prices, guarantees, regulated advice, and actions it can't take. Turn the high-stakes ones into handoff triggers.

Can it handle more than one language? Yes — Agent Vani answers in Hindi, English, and Hinglish and follows the caller's lead. Set your default in the persona. See multilingual phone support.

How do I know the script actually works? Call it and try to break it — interrupt, get vague, go off-script, correct a detail, ask for a human. Fix what you hear on real calls, not what you imagined on paper.

Where to go next

Related

Try Agent Vani on your lines

Configure persona and numbers, then dial a test call. No card required to start the trial.

Start free trial