OpenAI Function Calling vs Anthropic Tool Use: Complete 2026 Guide

Tool use is how LLMs stop being chat toys and start being useful parts of a real application. Both OpenAI and Anthropic ship mature tool-calling APIs in 2026, but they are not identical and the differences matter when you build. This article compares the two on API shape, developer ergonomics, parallel execution, and production patterns, with working code for each.

What tool use actually does

You define a set of tools (functions, API calls, database queries) and describe each one to the model with a name, a description, and a JSON schema for its arguments. The model decides, based on the user’s message, whether to call one of your tools. If yes, it returns a structured message containing the tool name and the arguments. Your code executes the tool, sends the result back to the model, and the conversation continues.

Both providers follow this same mental model. The differences are in how you declare tools, how parallel tool calls work, and what the response looks like.

OpenAI function calling in 2026

OpenAI’s tool API has gone through three major iterations since 2023. The current form (2026) uses the Chat Completions API with a tools parameter:

from openai import OpenAI
client = OpenAI()

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city.",
        "parameters": {
            "type": "object",
            "properties": {
                "city": {"type": "string", "description": "City name"},
                "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
            },
            "required": ["city"],
        },
    },
}]

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "What is the weather in Austin?"}],
    tools=tools,
)
tool_call = response.choices[0].message.tool_calls[0]
print(tool_call.function.name)       # get_weather
print(tool_call.function.arguments)  # {"city": "Austin", "unit": "fahrenheit"}

You then execute the tool, append the result as a message with role “tool”, and call the API again. The second call gives the model the tool output and lets it produce the final user-facing answer.

Anthropic tool use in 2026

Anthropic uses the Messages API with a tools parameter. The tool definition shape is similar but not identical:

from anthropic import Anthropic
client = Anthropic()

tools = [{
    "name": "get_weather",
    "description": "Get the current weather for a city.",
    "input_schema": {
        "type": "object",
        "properties": {
            "city": {"type": "string", "description": "City name"},
            "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
        },
        "required": ["city"],
    },
}]

response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=1024,
    tools=tools,
    messages=[{"role": "user", "content": "What is the weather in Austin?"}],
)
for block in response.content:
    if block.type == "tool_use":
        print(block.name)   # get_weather
        print(block.input)  # {"city": "Austin", "unit": "fahrenheit"}

Note the differences: no outer type: "function" wrapper, input_schema instead of parameters, and the response is a list of content blocks (text and tool_use) rather than a single message with tool_calls attached.

Parallel tool calls

Both providers support calling multiple tools in a single turn in 2026, but the ergonomics differ.

OpenAI returns a list of tool_calls in the response. You execute them in parallel (or sequentially) and append each result as its own tool message, tagged with the matching tool_call_id.

Anthropic returns multiple tool_use blocks in the content array. You execute each and send back tool_result blocks in a single user message, matching by tool_use_id.

Both work. OpenAI’s model is slightly cleaner because the list structure is explicit at the top level. Anthropic’s model fits better when a single turn mixes text and tool use (the model can say “Let me check the weather and the time” and then emit two tool_use blocks in one response).

Error recovery

When a tool call fails (bad arguments, downstream error), both providers let you send back an error message and give the model a chance to recover. Anthropic’s error recovery is noticeably better in 2026: Claude tends to adjust its arguments and retry inside the same turn, while GPT-4o is more likely to escalate to a text response explaining the failure. For agent workflows that need resilience, Claude’s behavior saves round trips.

Streaming and tool calls

Both providers stream tool calls as they generate. OpenAI’s streaming emits deltas for each argument as it is being formed; you cannot act on the tool call until the arguments are complete. Anthropic streams tool_use blocks the same way. Neither streams tool execution itself (that is your code).

In practice, streaming is useful for the model’s text response but not for tool arguments. Design your UI to show the final tool call after it is complete, then stream the model’s post-tool response.

Which to pick for your use case

  • Simple assistant with one or two tools: either works. Pick whichever provider you already use.
  • Agent system with heavy tool chaining: Anthropic. Error recovery and single-turn tool fluency help.
  • Codebase with existing OpenAI SDK investment: OpenAI. The migration cost is not worth it.
  • Mixed text and tool output in one turn: Anthropic. The content-block model fits naturally.
  • Parallel tool execution with explicit coordination: OpenAI. The top-level tool_calls list is slightly cleaner.

Provider-agnostic patterns

If you want to support both providers behind one interface, wrap tool definitions in your own schema and transform at call time:

def to_openai_tool(my_tool):
    return {"type": "function", "function": {
        "name": my_tool["name"],
        "description": my_tool["description"],
        "parameters": my_tool["schema"],
    }}

def to_anthropic_tool(my_tool):
    return {
        "name": my_tool["name"],
        "description": my_tool["description"],
        "input_schema": my_tool["schema"],
    }

Keep your canonical tool definitions in one shape; transform them per-provider at the boundary. LangChain and LangGraph handle this automatically if you use their tool abstractions.

Frequently Asked Questions

Which has better tool-use accuracy, OpenAI or Anthropic?

Both are strong in 2026. Independent benchmarks (Berkeley Function Calling Leaderboard, Tool Use Benchmark) show Claude Opus 5.5 and GPT-5 trading first place depending on the task type. For simple single-tool calls both are above 95% accuracy. For multi-step tool chains, Claude pulls ahead slightly on error-recovery scenarios.

Can I use the same tool definitions with both providers?

Not directly. The outer shape differs (OpenAI nests under function, Anthropic does not; parameters vs input_schema). Keep canonical definitions in your own shape and transform at call time. LangChain and LangGraph do this for you.

Does Gemini support tool use the same way?

Yes, Google Gemini 2.5 Pro and Flash both support function declarations through the Gemini API and Vertex AI SDK. The shape is closer to OpenAI’s than Anthropic’s. If you build provider-agnostic infrastructure, Gemini fits the same pattern as OpenAI with minor key renames.

How do I handle tools that take a long time to execute?

Return the tool result when the execution finishes, which may take seconds. For truly long-running operations (minutes or more), return a “queued” tool result immediately with a job ID, then follow up in a later turn when the real result is ready. Neither OpenAI nor Anthropic blocks waiting for tool execution; you control the timing.

Can I call a tool with structured output (not just dict)?

Yes. The tool’s return value should be a JSON-serializable object (dict, list, string). Both providers include the entire returned object in the next prompt as context. If you return a complex object, format it as JSON before sending back to the model so it can parse the fields reliably.

Which should a new project use in 2026?

If you have no provider preference, lean toward Anthropic for agent-heavy work and OpenAI for everything else (chat assistants, content generation, classification). Either way, design your code so swapping providers is a two-day refactor, not a rewrite. The frontier models keep moving; locking into one provider rarely pays.

Leave a Comment