Ch. 30 · AI & LLM Engineering

LLM Function Calling and Agent Loops Explained

How LLM tool use works: schemas, the call-execute-return loop, validation, error results, step limits and when an agent beats a workflow.

~8 min readintermediateupdated Oct 6, 2026

“How does function calling work, and how would you build an agent that uses tools?” is now a standard question for backend and AI engineering roles. The weak answer is “the model calls the API”. The strong answer is precise about who does what: the model emits a structured request, your code decides whether to run it, runs it, and sends the result back. Interviewers follow up with failure cases (bad arguments, failing tools, loops, duplicate side effects) because that is where real agents break, and with design judgement: when an agent is worth its cost and when a fixed workflow is better.

Before you start

You should be comfortable with Python dictionaries, JSON and exceptions, and know roughly what a JSON Schema looks like. Knowing what idempotency means for an API helps for the guardrails section. The code runs on Python 3.12+ with the standard library. A scripted fake model replaces the real LLM, so you can see exactly which messages flow through the loop without calling a paid API. The message shapes are simplified; real APIs use similar fields under provider-specific names.

The short answer

You send the model a list of tools, each with a name, a description and a JSON Schema for its arguments. When the model decides a tool would help, it stops generating text and returns a tool call: the tool name plus JSON arguments (some APIs signal this with a stop reason such as tool_use or a finish reason such as tool_calls). Your code validates the arguments, executes the function, and appends the tool result, linked to the call’s id, to the conversation. Then you call the model again. An agent is this loop running until the model answers in plain text or a limit is reached. The model never executes anything; every side effect passes through your code, which is where validation, permissions, limits and logging belong.

How it works

A tool definition is documentation for the model. The description tells it when to use the tool, and the schema tells it how:

import json

TOOLS = {
    "get_order": {
        "description": "Look up an order by id. Returns status and total.",
        "parameters": {
            "type": "object",
            "properties": {"order_id": {"type": "string"}},
            "required": ["order_id"],
            "additionalProperties": False,
        },
    },
    "refund_order": {
        "description": "Refund a delivered order. Irreversible.",
        "parameters": {
            "type": "object",
            "properties": {"order_id": {"type": "string"}, "reason": {"type": "string"}},
            "required": ["order_id", "reason"],
            "additionalProperties": False,
        },
    },
}

ORDERS = {"A-100": {"status": "delivered", "total": 42.0}}
REFUNDED = set()

def get_order(order_id):
    if order_id not in ORDERS:
        raise LookupError(f"order {order_id} not found")
    return ORDERS[order_id]

def refund_order(order_id, reason):
    if order_id in REFUNDED:                       # idempotent: a retry must not refund twice
        return {"refunded": False, "note": "already refunded"}
    REFUNDED.add(order_id)
    return {"refunded": True, "amount": ORDERS[order_id]["total"]}

HANDLERS = {"get_order": get_order, "refund_order": refund_order}
python

The model was trained to produce calls that match the schema, but “trained to” is not “guaranteed to”. Some providers offer a strict mode that constrains generation to the schema; even then, your code must check business rules the schema cannot express, such as whether this user owns this order.

Step-by-step walkthrough

Step 1: Validate arguments before executing

TYPES = {"string": str, "number": (int, float), "integer": int, "boolean": bool}

def validate(schema, args):
    props, problems = schema["properties"], []
    for key in schema["required"]:
        if key not in args:
            problems.append(f"missing {key}")
    for key, value in args.items():
        if key not in props:
            problems.append(f"unexpected {key}")
        elif not isinstance(value, TYPES[props[key]["type"]]):
            problems.append(f"{key} must be {props[key]['type']}")
    return "; ".join(problems)
python

In production, use a real JSON Schema validator or typed models such as Pydantic; this hand-rolled version shows what is being checked.

Step 2: Turn every outcome into a tool result

def run_tool(call):
    def result(content, is_error):
        return {"tool_call_id": call["id"], "is_error": is_error, "content": content}
    if call["name"] not in HANDLERS:
        return result(f"unknown tool {call['name']}", True)
    try:
        args = json.loads(call["arguments"])
    except json.JSONDecodeError:
        return result("arguments are not valid JSON", True)
    problem = validate(TOOLS[call["name"]]["parameters"], args)
    if problem:
        return result(problem, True)
    try:
        return result(json.dumps(HANDLERS[call["name"]](**args)), False)
    except Exception as exc:                        # report the failure, don't crash the loop
        return result(str(exc), True)
python

Unknown tools, malformed JSON, schema violations and runtime exceptions all become error results the model can read. Models are good at correcting a call when told exactly what was wrong; they cannot correct a crash or an empty string.

Step 3: Write the loop with a step limit

def agent(model, user_message, max_steps=6):
    messages = [{"role": "user", "content": user_message}]
    for step in range(max_steps):
        reply = model(messages)
        messages.append(reply)
        if not reply.get("tool_calls"):
            return reply["content"], step + 1
        results = [run_tool(c) for c in reply["tool_calls"]]   # parallel calls: return all results together
        messages.append({"role": "tool", "results": results})
    return "Stopped: step limit reached, handing off to a human.", max_steps
python

The loop ends in exactly two ways: the model answers in text, or the step budget runs out. A production loop also enforces a token or cost budget and a wall-clock timeout, and logs every call and result for debugging.

Step 4: Run it with a scripted model

def scripted_model(script):
    replies = iter(script)
    def model(messages):
        if messages[-1]["role"] == "tool":
            for r in messages[-1]["results"]:
                print("   tool result:", "ERROR" if r["is_error"] else "ok", r["content"])
        return next(replies)
    return model

script = [
    {"role": "assistant", "tool_calls": [{"id": "c1", "name": "get_order", "arguments": '{"order_id": 100}'}]},
    {"role": "assistant", "tool_calls": [{"id": "c2", "name": "get_order", "arguments": '{"order_id": "A-100"}'}]},
    {"role": "assistant", "tool_calls": [{"id": "c3", "name": "refund_order",
                                          "arguments": '{"order_id": "A-100", "reason": "damaged"}'}]},
    {"role": "assistant", "content": "Your order A-100 was refunded: 42.00."},
]
print(agent(scripted_model(script), "My order A-100 arrived damaged, please refund it."))
#    tool result: ERROR order_id must be string
#    tool result: ok {"status": "delivered", "total": 42.0}
#    tool result: ok {"refunded": true, "amount": 42.0}
# ('Your order A-100 was refunded: 42.00.', 4)
python

The first call passes a number where a string is required. The validator rejects it, the model sees the reason and retries correctly. That self-correction is the main reason to return errors rather than raise them.

Worked scenario

A support team’s first agent had none of these guardrails. Tool exceptions were caught and returned as an empty string, the loop had no step limit, and refund_order simply issued a refund on every call.

Two incidents followed. A customer asked about order “B-1”, which did not exist. The lookup failed silently, the model saw an empty result, tried again, and kept trying until the request hit the provider’s context limit; each conversation like this cost dozens of model calls. Separately, the payment provider timed out after successfully processing a refund. The orchestration code retried the model turn, the model requested the refund again, and the customer was refunded twice.

The fixes map one to one onto the code above. Failures became explicit error results (“order B-1 not found”), which lets the model tell the user the order does not exist. The loop got a step cap and a handoff message:

looping = [{"role": "assistant", "tool_calls": [{"id": f"x{i}", "name": "get_order",
                                                 "arguments": '{"order_id": "B-1"}'}]} for i in range(10)]
print(agent(scripted_model(looping), "Where is order B-1?", max_steps=3))
#    tool result: ERROR order B-1 not found
#    tool result: ERROR order B-1 not found
# ('Stopped: step limit reached, handing off to a human.', 3)
python

And the refund became idempotent, keyed by order id (in a real system, by an idempotency key passed to the payment provider), so a retried call returns “already refunded” instead of moving money twice. The team also moved refunds above a threshold behind a human approval step.

Common mistake

  • “The model calls the API.” The model emits text in a structured format; your code makes every real call.
  • Trusting arguments because they match the schema. Schema-valid does not mean authorized. Check ownership, limits and permissions in code.
  • Raising exceptions out of the loop or returning empty results. Return clear error results so the model can recover or explain.
  • No limits. Without step, token and time budgets, one confused conversation can run up a large bill.
  • Vague tool descriptions and overlapping tools. The model chooses tools from their names and descriptions; ambiguous ones cause wrong calls. Fewer, well-described tools beat many similar ones.
  • Using an agent when a workflow would do. If the steps are known in advance, a fixed sequence of model calls is cheaper, faster and easier to test.

Verify the behavior

Test the harness without any model, using scripted replies, so failures are deterministic:

def test_bad_arguments_return_an_error_result():
    r = run_tool({"id": "t1", "name": "get_order", "arguments": '{"order_id": 7}'})
    assert r["is_error"] and "must be string" in r["content"]

def test_refund_is_idempotent():
    REFUNDED.clear()
    first = run_tool({"id": "t2", "name": "refund_order", "arguments": '{"order_id": "A-100", "reason": "x"}'})
    second = run_tool({"id": "t3", "name": "refund_order", "arguments": '{"order_id": "A-100", "reason": "x"}'})
    assert json.loads(first["content"])["refunded"] and not json.loads(second["content"])["refunded"]

def test_loop_stops_at_the_step_limit():
    calls = [{"role": "assistant", "tool_calls": [{"id": str(i), "name": "get_order",
              "arguments": '{"order_id": "B-1"}'}]} for i in range(5)]
    assert agent(scripted_model(calls), "x", max_steps=2)[1] == 2

test_bad_arguments_return_an_error_result(); test_refund_is_idempotent(); test_loop_stops_at_the_step_limit()
print("ok")
python

Then add end-to-end evals with the real model: given a request, did it call the right tools with the right arguments, in an acceptable number of steps?

Follow-up questions

Workflow or agent? Use a workflow (prompt chaining, routing, parallel calls, a fixed sequence) when the steps are predictable. Use an agent when the number and order of steps depend on what the model discovers, and the task is valuable enough to justify extra latency, cost and risk.

What are parallel tool calls? One model turn can request several independent calls. Run them concurrently, and return all their results together in the next message.

What is MCP? The Model Context Protocol is an open standard for exposing tools, resources and prompts to LLM applications through a common client-server interface, so a tool written once can be used by many hosts.

How do you test agents? Unit-test tools and the harness deterministically, then run scenario evals that check final outcomes and the tool-call trajectory, several trials each because runs vary.

Interview exercise

You are building an agent that can read a user’s calendar and send meeting invites. In testing, it sometimes sends an invite to the wrong person with a similar name, and once it sent the same invite three times after a network error. Propose changes to the tools and the loop.

Answer and reasoning

The wrong-person problem is an ambiguity problem, so remove the ambiguity from the tool interface. Split the action into find_contacts(name), which returns candidates with email addresses, and send_invite(event_id, attendee_emails), which accepts only exact addresses. If find_contacts returns more than one match, the agent must ask the user rather than pick. Sending is a side effect on other people, so require user confirmation by showing the drafted invite and waiting for an explicit yes. For duplicates, make send_invite idempotent: compute an idempotency key from the event and attendee list, store it, and return “already sent” on repeats; also distinguish retryable network errors in the harness from the model’s decision to call the tool again. Finally, cap steps, log every call, and add these two incidents as eval scenarios. The reasoning: design tools so that the wrong action is hard to express, and enforce safety in code, not in the prompt.

Continue learning

More in AI & LLM Engineering

esc