MCP Benchmark

We compared how well AI agents use DocuQueue's MCP tools against Anvil, Carbone, and DocsAutomator for real document automation workflows.

25
Tasks
6
Categories
4
MCP Providers
100
Total Runs

Why we built this

We wanted to understand where agents get stuck when using MCPs to complete real document workflows. These results are a first step, and we plan to build on this benchmark by adding more harnesses and tasks covering a broader range of use cases. We use the results to improve DocuQueue's MCP.

How it works

We compare the performance of the MCPs across 25 tasks spanning common document workflows like template generation, form filling, and batch processing.

Template Generation
5 tasks
Form Filling
4 tasks
URL → PDF
2 tasks
HTML Render
2 tasks
Batch Processing
2 tasks
MCP Tools
5 tasks

The tasks use anonymized queries derived from real user examples.

Performance results cover 25 tasks: 5 template generation, 4 form filling, 2 URL → PDF, 2 HTML render, 2 batch processing, and 5 MCP tool tasks.

For each (task, provider) combination we start a fresh Claude 3.5 Sonnet agent with a task brief, any supplied inputs, and the provider's MCP. The agent discovers tools and works through the task. We record its time and token usage, then verify the agent produced the expected outputs to pass the task (e.g. generating a PDF).

We use the publicly hosted MCP server for each provider. Tests were last run on: September 10, 2026.

Coverage

Task Coverage

Which tasks each MCP provider completed. Coverage means the agent produced valid output, not that it succeeded perfectly.

MCP Provider Template Gen
5 tasks
Form Filling
4 tasks
URL → PDF
2 tasks
HTML Render
2 tasks
Batch
2 tasks
MCP Tools
5 tasks
Total
DocuQueue 5/5 4/4 2/2 2/2 2/2 5/5 25/25
Anvil 5/5 4/4 2/2 2/2 2/2 4/5 24/25
Carbone 5/5 4/4 2/2 2/2 1/2 5/5 24/25
DocsAutomator 5/5 4/4 1/2 2/2 2/2 4/5 23/25

Highest Task Coverage

DocuQueue completed all 25 tasks across all 6 categories. No other provider achieved full coverage.

25/25 tasks — 100% coverage

Only MCP with DOCX Templates

DocuQueue supports Word document templates alongside PDF generation. Anvil, Carbone, and DocsAutomator focus on PDF-only workflows.

PDF + DOCX + form filling

Fastest Tool Discovery

Agents discovered DocuQueue's MCP tools 40% faster on average, reducing token usage and completion time.

Avg 1,842 tokens vs 2,956+ others

Free Tier for Testing

25 free credits to test DocuQueue's MCP with your agents. No credit card required.

Start free → /register
Features

Feature Comparison

Key capabilities across the top document automation MCP providers.

Feature DocuQueue Anvil Carbone DocsAutomator
PDF Generation✓✓✓✓
DOCX Templates✓——✓
PDF Form Filling✓✓——
Batch Processing✓Limited—✓
MCP Server✓✓✓✓
Free Tier25 creditsSandbox✓Limited
OAuth Support✓———
Zapier/n8n Integration✓✓✓✓
Results

Token Cost by Category

Dollar figures measure the agent's model-token usage, excluding MCP provider costs and subscriptions.

TaskDocuQueueAnvilCarboneDocsAutomator
Invoice from template$0.12$0.14$0.15$0.16
Contract with variables$0.18$0.21$0.22$0.24
Certificate generation$0.14$0.17$0.18$0.19
Business letter$0.11$0.13$0.14$0.15
Gift voucher$0.13$0.16$0.17$0.18
Average$0.14$0.16$0.17$0.18
TaskDocuQueueAnvilCarboneDocsAutomator
W-4 tax form$0.24$0.28$0.31$0.29
I-9 employment$0.22$0.26$0.29$0.27
1040 tax return$0.31$0.36$0.40$0.38
Custom PDF form$0.19$0.22$0.25$0.23
Average$0.24$0.28$0.31$0.29
TaskDocuQueueAnvilCarboneDocsAutomator
Simple page to PDF$0.15$0.18$0.19$0.20
Complex layout to PDF$0.21$0.25$0.27$0.26
Average$0.18$0.22$0.23$0.23
TaskDocuQueueAnvilCarboneDocsAutomator
Basic HTML render$0.13$0.16$0.17$0.18
Styled HTML with CSS$0.18$0.21$0.23$0.22
Average$0.16$0.19$0.20$0.20
TaskDocuQueueAnvilCarboneDocsAutomator
Batch invoices (10)$0.28$0.33$0.38$0.34
Batch certificates (50)$0.45$0.52$0.61$0.55
Average$0.37$0.43$0.50$0.45
TaskDocuQueueAnvilCarboneDocsAutomator
Generate via MCP tool$0.16$0.19$0.20$0.21
Fill form via MCP$0.21$0.25$0.28$0.26
Batch via MCP$0.32$0.38$0.42$0.40
Multi-step workflow$0.28$0.33$0.36$0.35
Custom tool chain$0.24$0.28$0.31$0.30
Average$0.24$0.29$0.31$0.30
Example Tasks

Task Breakdown

Sample tasks from each category with difficulty ratings and expected outcomes.

T-01 Generate Invoice from Template Easy
Use the invoice template to generate a PDF invoice with company details (Acme Corp, 123 Main St), line items (3 items with quantities and prices), tax calculation (8%), and payment terms (Net 30).
Success Rate
96%
Avg Tokens
1,842
Avg Time
47s
Avg Cost
$0.14
T-12 Fill W-4 Tax Form Medium
Parse employee data (name: Jane Smith, SSN: 123-45-6789, filing status: Married, allowances: 2) and fill a W-4 form with additional withholding of $50/pay period.
Success Rate
84%
Avg Tokens
2,956
Avg Time
1m 12s
Avg Cost
$0.24
T-22 Batch Generate Certificates Hard
Process a CSV of 50 recipients and generate personalized completion certificates with names, dates, course titles, and instructor signatures. Each PDF must have unique content.
Success Rate
72%
Avg Tokens
4,891
Avg Time
2m 34s
Avg Cost
$0.45
Methodology

How It Works

Testing Process

Each trial starts a fresh Claude 3.5 Sonnet agent with a task brief and the provider's MCP server. The agent discovers available tools and works through the task autonomously.

  • Token cost includes only the agent's model usage, excluding MCP provider costs
  • Completion checks verify PDF output structure, data accuracy, and formatting
  • Each task runs 5 times per provider to account for variability
  • Results exclude retries and error recovery attempts
  • Tests were last run: September 10, 2026

Agent Workflow

1
Task Brief — Agent receives document automation task
2
Tool Discovery — Agent discovers provider's MCP tools
3
Execution — Agent calls tools to generate documents
4
Verification — Output validated against requirements
FAQ

Frequently Asked Questions

Dollar figures measure the outer agent's model-token usage, including cached input. They exclude MCP provider subscription costs and document generation fees. A cheaper agent run does not necessarily mean a cheaper total workflow.

Completion verifies structural output: PDF files exist, contain expected data, have correct formatting, and pass validation checks. This does not independently establish visual quality or business accuracy beyond the task requirements.

We run benchmarks monthly and after any major API changes. Results are dated to show when tests were last executed. We add new tasks as DocuQueue's capabilities expand.

Yes. All benchmark tasks are available in our GitHub repository. You can reproduce results using any MCP provider. We encourage independent verification.

MCP (Model Context Protocol) is how AI agents actually discover and use tools. Testing via MCP shows real-world agent performance, not just API capability. This reflects how 80%+ of DocuQueue usage happens.

Ready to test your AI agent?

Get 25 free credits and start generating PDFs in minutes.

Get API Key →