Poking at LLM benchmarks, 1 - Automation Bench
First of a few posts looking at the benchmarks used to evaluate LLMs.
Each new model release decimates the older benchmarks but outside models from the Openai/Anthropic duopoly and Deepseek, I still don't find the others reliable enough to be a daily driver. Would be interesting to see which benchmarks are actually a good proxy for things I use the models for -interactive coding (no software factories here).
I wanted to pick benchmarks that aren't saturated and yet something smaller models can make a reasonable pass at. The plan is to run a handful of tasks from each with a cheap model, read the trajectories, and see how the models go about solving them.
- We haven't got the budget for wasteful runs with frontier models, and the non open source ones don't expose thinking traces either which defeats the whole purpose (yep, I know chain-of-thought isn't causal).
- Additionally, benchmarks like Terminal bench also require a lot of background knowledge, which smaller models aren't going to be great at.
Newer model releases (Claude Fable, GPT Astra) reported scores on AutomationBench - a benchmark created by Zapier, which focuses on business process workflows. It includes regular processes enterprises set up automations for - reconciling excel transactions, or setting up onboarding for new hires (Digital paperwork?).
This feels like a good fit for our objective:
- External knowledge required is limited to knowing the popular software services, and useful know-how for the task lives in the environment. The model just has to find it, and string together tool calls with some judgement. That's trainable for small models.
- Frontier models still sit below 50% accuracy on it, so there's headroom to improve.
All the public tasks in AutomationBench are listed here.
How AutomationBench tasks were constructed
The arxiv paper covers the shape of the benchmark:
- The public task-set has 600 entries across different domains (support, financial, etc.), plus a list of 200 simple tasks.
- Completing a task requires the use of multiple (simulated) SaaS apps, out of a total of 47 (Google drive, Salesforce, Buffer, etc.). Each app provides a set of tools the agent can call.
- Each task specifies a list of assertions against which the final state is checked, with a strict pass/fail on each.
The tasks are LLM-generated1 by mining Zapier's usage data, and then "hardened" by adding decoy data entries with similar names, obscuring key information and having strict policy rules. Only a sub-sample of the tasks underwent manual inspection.
NOTE: The leaderboard on Zapier's website evaluates models on a held-out private task set, which was created from a similar distribution, but made harder.
Running the benchmark
The public taskset is available as a Github repo, with an evaluation harness and scorer built in. Agents can make upto 50 tool calls to complete the task at hand.
We used the Inkling-small model from Thinking Machines, via OpenRouter. It's fairly agentic, and less likely to be bench-maxxed per my uninformed guess (their business model is finetuning).
Evaluation can be run in one of three modes, called toolsets.
api: tools are available as REST API endpoints. The LLM only sees two tools - search, and execute.zapier: same as theapimode, but tools are regular python methods, no request wrangling necessary.limited_zapier: only the exact tools specified for a task are made available to it.
The limited_zapier mode kinda leaks information how to solve a task, by massively reducing the action space for the model! For real world business automations, this is something the AI agent must figure out from the instructions.
We started by running 15 randomly sampled tasks, with the zapier toolset and the model at High thinking level.
Inkling-smallonly got 2 of them fully right (all assertions passed).- On some tasks, it couldn't even get started. Whereas for others, it managed partial credit - updating state that wasn't needed, missing the right records, running over the tool calls limit.
However, reading through the traces pointed out that 6 or 7 tasks had issues that either makes the task unsolvable, or let wrong answers through. Which feels like a big indictment on the quality of the benchmark.
Issues with the public task set
Some of the bugs are listed below. I filed a concrete github issue at zapier/AutomationBench#24.
Checks that aren't strict enough
- Task: Finance, Subscription billing.
- The agent needs to check customer subscriptions up for renewal, and create invoices for them with the updated rate.
- The rubric includes an assertion that a churned customer's record isn't updated. This is checked by looking for an
updatedflag set on the corresponding row. The spreadhseet tool to update cells makes sure to set that flag. - However, this misses another way sheets can be updated - deleting the original row, and recreating it with the new values. Which is exactly what the model did in one of the traces, and wasn't penalized for.
- The issue doesn't seem limited to this task, since there are 42 tasks (97 assertions) which rely on the same verifier method
google_sheets_row_not_updated.
Tool errors that look like data
- Task: Marketing, Ad performance review.
- Agent needs to scan for the active campaigns, and send an email with list of all that were paused for underperformance.
- The model tried to call the method
gmail_send_emaildirectly, instead of passing the method name and arguments to theexecute_tooltool. - The runtime environment catches every exception and returns
str(e)as the tool result. For a KeyError that's the quoted key. Since there are only two top level tools, the agent gets back"gmail_send_email"in the response.
Without any error markers, the agent concludes this tool call as successful and moves on.
- Since no actual email records got created, it fails the task; whereas with a clear error message, it could have corrected.
- Ideally, the harness returns a clear response indicating that a tool call failed. Something like
{"error": ..., "detail": ...}, or{"success": False, ...}. - This is also a general issue, not limited to this task. It is likely that stronger models don't make the same mistake and enver run into it but worth fixing anyway.
Incomplete context supplied to agent
- Task: Finance, Late fee calculation
- Agent needs to look for outstanding invoices, and apply late fees to those that are past due. Late fee calculation depends on the count of days overdue. The rubric checks the expected amounts for each record, as computed from
meta.current_time. - However, the evaluation harness doesn't include the metadata property in the context supplied to the model! And there is no
clocktool it can leverage to fetch it. - The model ended up guessing the date from an inbox email, computed a different fee, and failed. Frontier models will fail too unless they can intuit the correct date.
- Almost half the tasks set a currrent timestamp in the metadata, and around 50 of them use relative-time wording ("today", "overdue", "N days") despite not providing the date in the prompt.
Incorrect assertion in rubric
- Task: Sales, Recency selection
- Agent is tasked to update the contact details for a customer, also noting the source email request. Not particularly challenging, and the model completed the task successfully.
- The rubric contains an assertion checking that id of the source email is present in the updated record. However, another assertion mistakenly checks for presence of another email's id.
- The second assertion should be a "not-exists", rather than "exists" check. Feels like a typo during task generation.
Unsolvable task
- Task: Finance, New hire onboarding
- The prompt asks the agent to refer to a setup checklist, and execute all the steps mentioned there, without naming the file directly.
- The data lives in a Google sheets spreadsheet. However, in the
limited_zapiermode, the list of tools specified doesn't include any tools that let you search over files based on their name/contents. - So, the model resorts to inventing names and try see if they exist.
Prompt with insufficient context
- Task: Marketing, Video repurpose
- Agent is asked to go over the video library and based on some criteria, pick videos to create clips from and "queue" them. Nothing is specified about the output destination.
- The agent wrote the final output to
Buffer, because Buffer has an "add to queue" action. However, the rubric was looking for the agent to write to a spreadsheet calledss_clips. - The current setup is testing whether the model can intuit organizational preferences about what data lives where, which is kinda unfair.
Takeaways
TLDR: Good shape, but not a well-vetted benchmark. Hard to trust the leaderboard.
AutomationBench is set up well. Zapier is a really good data source for what actually deployed business workflows are like. You do lose some fidelity moving from the day-to-day environment to a static evaluation - users' intent evolves, available services change, and more. Still, this feels much closer to real world automation usage than creating GCC from scratch.
However, the task/rubric quality isn't great. Looking over agent traces only from a small sample of tasks surfaced issues in half of them. Unless the private task set underwent much more diligence, hard to trust the results coming out of it.
Some other notes.
-
User instructions are generally missing context on what information lives where, which is unrealistic. Typically, this kind of organizational know-how is available to employees and any deployed AI agents. Note that there is no list tool, which the agent can use to figure this out themselves. Including a summary that lists all available services, and the domains each is used for, would cut down a big chunk of tool calls agents have to make.
-
Some of the tasks like Subscription billing have policy notes in data ("don't delete rows") that contradict the user instructions ("remove churned customer records"). This is a good behavioral check - models (both Inkling-small and Muse-spark-1.3) sided with what was present in the document, which feels right.
-
Using a weak model for evaluation. I feel like using a smaller agentic model to run evaluation was very useful. Frontier models can just power through underspecified tasks, and are less likely to land in non-happy paths. So you can't really tell if the rubric is sufficiently discriminative.
-
Inspect is a neat tool to visualize agent traces. The raw views are a little hard to scan through, but can be easily customized with your favorite coding agent.
Footnotes
-
Models used for task generation were Opus 4.6, GPT 5.3 Codex, Gemini 3, so a little dated at this point. ↩