We compared how well AI agents use DocuQueue's MCP tools against Anvil, Carbone, and DocsAutomator for real document automation workflows.
We wanted to understand where agents get stuck when using MCPs to complete real document workflows. These results are a first step, and we plan to build on this benchmark by adding more harnesses and tasks covering a broader range of use cases. We use the results to improve DocuQueue's MCP.
We compare the performance of the MCPs across 25 tasks spanning common document workflows like template generation, form filling, and batch processing.
The tasks use anonymized queries derived from real user examples.
Performance results cover 25 tasks: 5 template generation, 4 form filling, 2 URL → PDF, 2 HTML render, 2 batch processing, and 5 MCP tool tasks.
For each (task, provider) combination we start a fresh Claude 3.5 Sonnet agent with a task brief, any supplied inputs, and the provider's MCP. The agent discovers tools and works through the task. We record its time and token usage, then verify the agent produced the expected outputs to pass the task (e.g. generating a PDF).
We use the publicly hosted MCP server for each provider. Tests were last run on: September 10, 2026.
Which tasks each MCP provider completed. Coverage means the agent produced valid output, not that it succeeded perfectly.
| MCP Provider | Template Gen 5 tasks |
Form Filling 4 tasks |
URL → PDF 2 tasks |
HTML Render 2 tasks |
Batch 2 tasks |
MCP Tools 5 tasks |
Total |
|---|---|---|---|---|---|---|---|
| DocuQueue | 5/5 | 4/4 | 2/2 | 2/2 | 2/2 | 5/5 | 25/25 |
| Anvil | 5/5 | 4/4 | 2/2 | 2/2 | 2/2 | 4/5 | 24/25 |
| Carbone | 5/5 | 4/4 | 2/2 | 2/2 | 1/2 | 5/5 | 24/25 |
| DocsAutomator | 5/5 | 4/4 | 1/2 | 2/2 | 2/2 | 4/5 | 23/25 |
DocuQueue completed all 25 tasks across all 6 categories. No other provider achieved full coverage.
DocuQueue supports Word document templates alongside PDF generation. Anvil, Carbone, and DocsAutomator focus on PDF-only workflows.
Agents discovered DocuQueue's MCP tools 40% faster on average, reducing token usage and completion time.
25 free credits to test DocuQueue's MCP with your agents. No credit card required.
Key capabilities across the top document automation MCP providers.
| Feature | DocuQueue | Anvil | Carbone | DocsAutomator |
|---|---|---|---|---|
| PDF Generation | ✓ | ✓ | ✓ | ✓ |
| DOCX Templates | ✓ | — | — | ✓ |
| PDF Form Filling | ✓ | ✓ | — | — |
| Batch Processing | ✓ | Limited | — | ✓ |
| MCP Server | ✓ | ✓ | ✓ | ✓ |
| Free Tier | 25 credits | Sandbox | ✓ | Limited |
| OAuth Support | ✓ | — | — | — |
| Zapier/n8n Integration | ✓ | ✓ | ✓ | ✓ |
Dollar figures measure the agent's model-token usage, excluding MCP provider costs and subscriptions.
| Task | DocuQueue | Anvil | Carbone | DocsAutomator |
|---|---|---|---|---|
| Invoice from template | $0.12 | $0.14 | $0.15 | $0.16 |
| Contract with variables | $0.18 | $0.21 | $0.22 | $0.24 |
| Certificate generation | $0.14 | $0.17 | $0.18 | $0.19 |
| Business letter | $0.11 | $0.13 | $0.14 | $0.15 |
| Gift voucher | $0.13 | $0.16 | $0.17 | $0.18 |
| Average | $0.14 | $0.16 | $0.17 | $0.18 |
| Task | DocuQueue | Anvil | Carbone | DocsAutomator |
|---|---|---|---|---|
| W-4 tax form | $0.24 | $0.28 | $0.31 | $0.29 |
| I-9 employment | $0.22 | $0.26 | $0.29 | $0.27 |
| 1040 tax return | $0.31 | $0.36 | $0.40 | $0.38 |
| Custom PDF form | $0.19 | $0.22 | $0.25 | $0.23 |
| Average | $0.24 | $0.28 | $0.31 | $0.29 |
| Task | DocuQueue | Anvil | Carbone | DocsAutomator |
|---|---|---|---|---|
| Simple page to PDF | $0.15 | $0.18 | $0.19 | $0.20 |
| Complex layout to PDF | $0.21 | $0.25 | $0.27 | $0.26 |
| Average | $0.18 | $0.22 | $0.23 | $0.23 |
| Task | DocuQueue | Anvil | Carbone | DocsAutomator |
|---|---|---|---|---|
| Basic HTML render | $0.13 | $0.16 | $0.17 | $0.18 |
| Styled HTML with CSS | $0.18 | $0.21 | $0.23 | $0.22 |
| Average | $0.16 | $0.19 | $0.20 | $0.20 |
| Task | DocuQueue | Anvil | Carbone | DocsAutomator |
|---|---|---|---|---|
| Batch invoices (10) | $0.28 | $0.33 | $0.38 | $0.34 |
| Batch certificates (50) | $0.45 | $0.52 | $0.61 | $0.55 |
| Average | $0.37 | $0.43 | $0.50 | $0.45 |
| Task | DocuQueue | Anvil | Carbone | DocsAutomator |
|---|---|---|---|---|
| Generate via MCP tool | $0.16 | $0.19 | $0.20 | $0.21 |
| Fill form via MCP | $0.21 | $0.25 | $0.28 | $0.26 |
| Batch via MCP | $0.32 | $0.38 | $0.42 | $0.40 |
| Multi-step workflow | $0.28 | $0.33 | $0.36 | $0.35 |
| Custom tool chain | $0.24 | $0.28 | $0.31 | $0.30 |
| Average | $0.24 | $0.29 | $0.31 | $0.30 |
Sample tasks from each category with difficulty ratings and expected outcomes.
Each trial starts a fresh Claude 3.5 Sonnet agent with a task brief and the provider's MCP server. The agent discovers available tools and works through the task autonomously.
Dollar figures measure the outer agent's model-token usage, including cached input. They exclude MCP provider subscription costs and document generation fees. A cheaper agent run does not necessarily mean a cheaper total workflow.
Completion verifies structural output: PDF files exist, contain expected data, have correct formatting, and pass validation checks. This does not independently establish visual quality or business accuracy beyond the task requirements.
We run benchmarks monthly and after any major API changes. Results are dated to show when tests were last executed. We add new tasks as DocuQueue's capabilities expand.
Yes. All benchmark tasks are available in our GitHub repository. You can reproduce results using any MCP provider. We encourage independent verification.
MCP (Model Context Protocol) is how AI agents actually discover and use tools. Testing via MCP shows real-world agent performance, not just API capability. This reflects how 80%+ of DocuQueue usage happens.