Tool calling in AI agents is the mechanism that lets a large language model request work from software outside the model. Instead of only generating a sentence, the model can produce a structured request such as get_order_status(orderId: "A1042"). Your application validates that request, runs the corresponding function or API, and sends the result back to the model.
That sounds simple, but it changes the role of an LLM significantly. A chat model can explain how to reschedule a delivery. An agent with the right tool can retrieve the delivery, check the authenticated user's permissions, find available dates, and request the change.
The distinction matters: the model proposes a tool call; your application executes it. The LLM should not hold database credentials, bypass business rules, or become the authorization layer.
What Is Tool Calling in AI Agents?
Tool calling—also called function calling—is a contract between an LLM and the application hosting it. The application describes one or more capabilities in a machine-readable form. Each description normally contains:
- a unique tool name;
- a precise description of when the tool should be used;
- an input schema, usually expressed with JSON Schema;
- sometimes an output schema or result contract;
- application-side code that performs the actual operation.
For example, an order-support agent might receive this simplified definition:
{
"type": "function",
"name": "get_order_status",
"description": "Return the current status of an order owned by the authenticated customer.",
"parameters": {
"type": "object",
"properties": {
"order_id": {
"type": "string",
"description": "The order identifier, for example A1042."
}
},
"required": ["order_id"],
"additionalProperties": false
},
"strict": true
}
If the user asks, “Where is order A1042?”, the model may return a tool-call item containing the function name and arguments instead of a final natural-language answer. The host application then performs the lookup and submits the tool result to the model. Only after seeing that result does the model produce a response such as, “Order A1042 left the distribution center this morning and is expected tomorrow.”
Tool calling does not give the model arbitrary access to your code. It gives the model a controlled vocabulary of actions that your application has chosen to expose.
Why normal chat is not enough
An LLM generates output from the context supplied to it and patterns learned during training. Without external capabilities, it has several important limitations.
It does not automatically know live or private data
A model cannot know the current status of a customer's order, today's inventory, a private account balance, or an incident created five minutes ago unless the application supplies that information. Asking it to “answer confidently” does not solve the data-access problem; it increases the chance of a convincing fabrication.
Text is not an authenticated business action
A message saying “Your appointment has been moved” does not update a calendar. A real change requires authenticated API access, validation, concurrency handling, persistence, and a confirmed result.
Business operations need deterministic boundaries
Production systems need typed inputs, explicit permissions, timeouts, retry policies, audit records, and predictable failure behavior. Free-form text is a poor interface for those requirements. A tool schema narrows the model's request into something ordinary application code can inspect.
Normal chat is still appropriate when the task is explanation, summarization, drafting, or reasoning entirely over supplied context. Tool calling becomes useful when the answer depends on a computation, a current lookup, a private system, or a real-world action.
How an LLM decides to call a tool
At request time, the model receives the conversation, developer instructions, and the tools available for that turn. It evaluates whether a tool is relevant and, if so, generates a structured tool request. Depending on the API, the application can usually allow automatic selection, require a particular tool, restrict the eligible subset, or prevent tool use.
The decision is influenced by several inputs:
- The user's intent. “What is compound interest?” may need no tool. “Calculate the monthly payment for this loan” may justify a calculator.
- The tool name and description. Clear, distinct descriptions improve selection. Overlapping or vague tools make routing less reliable.
- The parameter schema. Field names, types, enums, and descriptions tell the model what arguments it must provide.
- System and developer instructions. These can define conditions such as requiring confirmation before a write.
- The available context. If an order ID is missing, the model should ask for it rather than invent it.
This is model inference, not a conventional if/else router. Even when a schema constrains the shape of the arguments, it does not prove that the requested action is correct, authorized, or wise. Tool selection should therefore be tested with evaluation datasets that include normal requests, ambiguous inputs, adversarial prompts, missing parameters, and requests that should not call any tool.
The tool-calling execution loop
A reliable implementation is a loop owned by the host application:
- The application sends the user's request and eligible tool definitions to the model.
- The model returns either a normal response or one or more tool-call requests.
- The application parses the request and validates it against the schema.
- The policy layer checks identity, authorization, scope, risk, and approval requirements.
- The application invokes the mapped function or service.
- The tool returns a typed success or failure result.
- The application associates the result with the original tool-call ID and sends it back to the model.
- The model may request another tool or produce the final user-facing answer.
That loop needs stopping conditions. Set limits for tool-call count, elapsed time, token use, retries, and repeated calls with identical arguments. Without them, a malformed result or poor tool description can create an expensive loop.
Function and API examples
A “tool” does not have to be a public HTTP endpoint. It can be any bounded operation the host knows how to execute.
| User request | Possible tool | What the application actually does |
|---|---|---|
| “What is the weather in Raleigh?” | get_weather | Calls a weather provider with server-held credentials. |
| “How much will this loan cost?” | calculate_payment | Runs deterministic financial code with validated numeric inputs. |
| “Where is my order?” | get_order_status | Queries an internal service using the authenticated customer's scope. |
| “Book the 3 p.m. appointment.” | book_appointment | Requires confirmation, checks availability again, writes the booking, and returns a receipt. |
| “Summarize our refund policy.” | search_policy_documents | Retrieves relevant approved document passages for a grounded answer. |
Good tools are narrow and business-oriented. Prefer cancel_order(order_id, reason) over execute_sql(query). Prefer create_support_ticket over a generic http_request. Narrow tools reduce the action space, make authorization understandable, and produce cleaner audit logs.
Tool calling vs MCP
Tool calling and the Model Context Protocol (MCP) solve related problems at different layers.
| Question | Tool calling | MCP |
|---|---|---|
| What is it? | A model interaction pattern for requesting a function with structured arguments. | An open protocol for connecting AI applications to servers that expose tools, resources, and prompts. |
| Main concern | Selection, arguments, invocation, result return, and continuation. | Discovery, transport, capability negotiation, and interoperable integration. |
| Who defines the connection? | Usually application-specific SDK and code. | MCP client and server follow a shared protocol. |
| Does it execute securely by itself? | No. | No. The host and server still need authentication, authorization, validation, and policy. |
An application can use tool calling without MCP by declaring local functions directly. It can also use MCP to discover tools from one or more servers and present the permitted tools to the model. MCP is therefore not a replacement for tool calling; it can be the standardized integration layer that supplies callable capabilities.
Tool calling vs RAG
Retrieval-augmented generation (RAG) retrieves relevant information and places it in the model's context so the model can produce a grounded answer. Tool calling is the broader mechanism for requesting an external operation.
| Dimension | RAG | Tool calling |
|---|---|---|
| Primary goal | Ground an answer in retrieved knowledge. | Obtain data, run computation, or perform an action. |
| Typical operation | Search, rank, and return document passages. | Invoke any exposed function or API. |
| Side effects | Usually read-only. | May be read-only or state-changing. |
| Output | Evidence or context for generation. | A typed result, error, receipt, or data payload. |
The two commonly work together. An agent may call a retrieval tool to find policy passages, answer with citations, and then—after explicit approval—call a separate tool to open a case. In that design, retrieval is one tool-backed capability within a larger workflow.
The connection to structured output
Tool calling and structured output both use schemas, but they have different purposes.
- Tool calling asks the model to select an operation and produce its arguments.
- Structured output asks the model to format its response as a predictable object for downstream code.
Suppose a claims assistant must decide whether to retrieve a policy. The tool call could be:
{
"name": "search_claims_policy",
"arguments": {
"state": "NC",
"topic": "rental vehicle coverage"
}
}
After the lookup, the final model response might use a separate structured-output schema:
{
"coverage_status": "requires_review",
"summary": "Coverage depends on the selected endorsement.",
"citations": ["policy-2026-04#section-8.2"],
"next_action": "route_to_adjuster"
}
Strict schema adherence reduces parsing failures, but schema-valid is not the same as business-valid. An order ID can be a valid string and still belong to another customer. A refund amount can be a valid number and still exceed the agent's authority. Treat schema validation as the first gate, not the last.
Security risks in tool-enabled agents
Adding tools changes an LLM application from a content-generation system into a potential action system. The risk is not merely that the model gives a wrong answer; it may request the wrong operation against a real service.
Prompt injection
Instructions can arrive through user input, retrieved documents, web pages, emails, or tool results. An untrusted document might contain text telling the model to ignore policy and export data. Treat external content as data, not authority. Do not allow retrieved text to redefine permissions.
Excessive agency and over-permissioned tools
A generic administrative API gives the model more capability than most tasks require. Apply least privilege to both the tool catalog and the credentials behind each tool. Expose only the tools needed for the current user, tenant, workflow stage, and conversation.
Broken authorization
Never trust identifiers supplied by the model as proof of access. Derive identity and tenant scope from the authenticated session. Recheck authorization inside the tool handler, just as you would for an ordinary controller or API endpoint.
Unsafe writes and duplicate actions
Retries and repeated model calls can create duplicate orders, tickets, transfers, or messages. Use idempotency keys, optimistic concurrency, transactional boundaries, and explicit confirmation for consequential actions. Separate “preview” from “commit” when possible.
Data leakage through arguments, results, and logs
Minimize sensitive fields sent to the model. Redact secrets and personal data from logs. Do not place service credentials in prompts or tool descriptions. Define retention and telemetry policies for model requests, tool inputs, outputs, and traces.
Denial of wallet and runaway loops
Set per-request budgets for model turns, parallel calls, result size, retries, and wall-clock duration. Cache safe read operations where appropriate and reject repeated calls that cannot make progress.
A production architecture for tool calling
A production design should keep probabilistic reasoning separate from deterministic enforcement.
User or client
↓
API gateway and authentication
↓
Agent orchestrator ──→ model gateway
↓ ↓
Tool registry tool-call proposal
↓ ↓
Policy and approval layer ←┘
↓
Typed tool adapter
↓
Domain service / API / database
↓
Sanitized tool result → orchestrator → model → final response
Cross-cutting: tracing, audit, evaluation, budgets, secrets, retries
The components have distinct responsibilities:
- API gateway: authenticates the caller, applies rate limits, and establishes tenant context.
- Agent orchestrator: owns conversation state, model calls, tool-call loops, budgets, and stopping rules.
- Tool registry: contains versioned names, descriptions, schemas, risk classification, and handler mappings.
- Policy layer: decides whether a proposed call is allowed, needs confirmation, or must be rejected.
- Tool adapter: converts validated arguments into a domain command or service request and normalizes errors.
- Domain service: remains the source of business rules and authorization. The agent does not bypass it.
- Observability: records correlation IDs, selected tools, latency, status, token usage, policy decisions, and sanitized results.
- Evaluation: measures tool-selection accuracy, argument correctness, unnecessary calls, refusal behavior, task completion, cost, and latency.
For high-impact actions, add a human approval checkpoint after the tool proposal and before execution. Show the user the exact action and material parameters—not merely “Allow?”—and bind the approval to that specific request.
Implementation direction for .NET, Python, and Go
The core design should stay the same across languages: define typed contracts, expose JSON schemas, map approved calls to ordinary application services, and keep the orchestration loop outside domain logic.
.NET
For a .NET system, model each tool as a small service method with strongly typed request and result records. Use ASP.NET Core dependency injection for domain clients, HttpClientFactory for outbound APIs, resilience policies for transient failures, and OpenTelemetry for traces. Microsoft Semantic Kernel can provide plugin discovery and automatic or manual function invocation; direct provider SDK calls are also reasonable when you want a thinner orchestration layer.
public sealed record GetOrderStatusArgs(string OrderId);
public interface IOrderTools
{
Task<OrderStatusResult> GetOrderStatusAsync(
GetOrderStatusArgs args,
UserContext user,
CancellationToken cancellationToken);
}
Prefer manual invocation for sensitive tools so your policy code can inspect the call before the handler runs. Keep authorization inside IOrderTools or the underlying domain service even if the orchestrator already checked it.
Python
Python is a strong choice for experimentation and AI-heavy services. Use Pydantic models for tool inputs and outputs, async HTTP clients for I/O, and a registry that maps tool names to callables. Frameworks can reduce boilerplate, but the authorization, approval, and retry rules should remain explicit application code.
class GetOrderStatusArgs(BaseModel):
order_id: str
async def get_order_status(
args: GetOrderStatusArgs,
user: UserContext
) -> OrderStatusResult:
await authorize_order_access(user, args.order_id)
return await order_client.get_status(args.order_id)
Generate the model-facing schema from the same typed definition when practical. This reduces drift between what the model is told and what the handler accepts.
Go
Go fits well when the orchestration service needs simple deployment, predictable concurrency, and a small runtime footprint. Define argument structs with JSON tags, validate them after unmarshalling, pass context.Context through every call, and use an explicit switch or typed registry to dispatch tools.
type GetOrderStatusArgs struct {
OrderID string `json:"order_id"`
}
func (s *ToolService) GetOrderStatus(
ctx context.Context,
user UserContext,
args GetOrderStatusArgs,
) (OrderStatusResult, error) {
if err := s.authorize(ctx, user, args.OrderID); err != nil {
return OrderStatusResult{}, err
}
return s.orders.GetStatus(ctx, args.OrderID)
}
Avoid burying tool dispatch in reflection-heavy abstractions until the tool set genuinely requires them. Explicit code is often easier to audit.
Practical design rules
- Give every tool one clear responsibility.
- Use precise descriptions and non-overlapping names.
- Use strict schemas where supported, then perform domain validation again.
- Expose the smallest tool set needed for the current task.
- Separate read tools from write tools and classify risk.
- Require explicit approval for consequential or irreversible operations.
- Use idempotency keys for state-changing calls.
- Return compact, typed results; do not dump entire database records into context.
- Give tools stable error codes the model can reason about safely.
- Trace the complete path from user request to tool call to final answer.
- Evaluate calls that should happen and calls that must not happen.
- Keep business rules in domain services, not in prompts.
Tool calling is the bridge—not the whole agent
Tool calling gives an LLM a structured way to request data and actions. It does not, by itself, provide memory, planning, retrieval quality, permissions, reliability, or observability. Those capabilities come from the surrounding application architecture.
This is why a useful mental model is: the LLM is a probabilistic planner and interface; the application is the trusted control plane. The model may decide that an order lookup is useful and propose the arguments. Your software decides whether the call is valid, permitted, affordable, and safe to execute.
If you are new to the broader architecture, start with What Is an AI Agent?, compare the boundaries in AI Agent vs Chatbot vs Copilot, and then read What Is Agentic AI?. Together, those concepts explain where tool calling fits: it is one of the main mechanisms that turns model output into controlled interaction with real systems.
Frequently asked questions
What is tool calling in an AI agent?
Tool calling is a structured interaction in which an LLM requests a named function and supplies arguments that match a declared schema. The host application validates and authorizes the request, executes the function, and returns the result to the model.
Does the LLM directly execute the function?
Usually, no. The model generates a tool-call request. The application or agent runtime maps that request to executable code. Some platforms provide hosted tools, but the security principle remains the same: the allowed tools and execution boundary are controlled by the host.
Is function calling the same as tool calling?
The terms are often used interchangeably. “Function calling” usually refers to developer-defined functions with structured arguments. “Tool calling” is broader and may include hosted search, file retrieval, code execution, computer use, or tools discovered through MCP.
How does the model know which tool to use?
The model uses the user's request, conversation context, system instructions, tool names, descriptions, and parameter schemas. Applications can also force, restrict, or disable tool selection. Because selection is probabilistic, it should be evaluated rather than assumed correct.
What is the difference between tool calling and MCP?
Tool calling is the model-to-function interaction pattern. MCP is a protocol that standardizes how AI applications connect to servers exposing tools, resources, and prompts. An MCP client can discover a server's tools and make them available for model tool calling.
What is the difference between tool calling and RAG?
RAG retrieves relevant content to ground a generated answer. Tool calling can perform retrieval, but it can also calculate values, query live systems, or change state. A retrieval operation is often implemented as one tool within an agent.
Does strict structured output make a tool call safe?
No. It can ensure that arguments match the expected schema, but it does not establish user authorization, ownership, business validity, or intent. Application-side policy and domain validation are still required.
Should an agent be allowed to call write APIs automatically?
Only when the risk is low and the action is reversible, well-scoped, authorized, and idempotent. Financial, legal, destructive, external-communication, or otherwise consequential actions should usually require an explicit approval step.
Internal linking suggestions for the editor
- From What Is an AI Agent?, link the phrase “agents use tools to take action” to this article.
- From What Is Agentic AI?, link the first discussion of action execution or external APIs to this article.
- From AI Agent vs Chatbot vs Copilot, link “ability to use tools” in the comparison section to this article.
- From What Is MCP?, link “tool calling” to this article and link this article's MCP comparison back to the MCP guide.
- When the planned RAG article is published, replace the provisional
/what-is-rag/link above with its canonical URL and add a reciprocal link from its “RAG in agents” section. - Future links: connect “strict schemas” to a Structured Outputs article, “evaluation datasets” to an LLM Evaluation article, and “prompt injection” to an AI Agent Security article.
