Spots

Launching Vespper DOCX MCP: 3× faster, 2× cheaper, more accurate

Word documents are everywhere. In domains such as legal, finance, and healthcare, the Word document is the deliverable. Contracts, regulatory submissions, and audit reports get drafted, redlined, and signed in .docx, and companies usually have libraries of Word templates they work with on a regular basis.

That work is increasingly shifting to agents. Microsoft

That work is increasingly shifting to agents. Microsoft Copilot and Claude in Word have made in-document AI mainstream, while a growing number of vertical agents; particularly in legal tech—need to interact with .docx files. However, AI agents are still struggling to do great work in Word documents.

We spoke with dozens of software engineers, mostly

We spoke with dozens of software engineers, mostly in the legal tech and health care spaces, who said they spend weeks and even months tuning their harnesses to edit Word documents reliably, and now are forced to maintain very complex in-house solutions. In general, there exist several ways for agents to edit Word docs today, which mainly fall into three categories: Letting agents write code that uses low-level SDKs such as python-docx / aspose / Open XML SDK. Connect MCPs such as SuperDoc, Office CLI, safe-docx or Adeu that provide opinionated tools for the agent.

Round-trip the Word document through a lossy projection

Round-trip the Word document through a lossy projection i.e. convert to Markdown/HTML using something like pandoc/mammoth.js, let the agent edit it and convert it back to .docx. The current solutions work on simple cases, but they fall short when it comes to complex scenarios. Before diving deep into the solutions and their drawbacks, let’s first understand what a .docx file is.

A .docx file is essentially a ZIP file

A .docx file is essentially a ZIP file with a hierarchy of XML files following the OOXML (Office Open XML) spec. Inside the ZIP, there are the following files: document.xml contains the main text, styles.xml defines reusable styles (kind of like a CSS stylesheet), numbering.xml defines list/numbering behavior, and separate XML files store headers, footers, footnotes, relationships, media, and document metadata.

These XML files are quite verbose. For example

These XML files are quite verbose. For example, in document.xml, even a short 4–5-sentence paragraph can turn into thousands of tokens once you add styles, metadata, formatting information, run splitting, and XML boilerplate. The text users see in Microsoft Word might be split across many XML nodes and might be persisted in a very verbose manner. Here's an interactive widget that shows how a simple .docx file works behind the scenes:

This nature of DOCX editing makes it different

This nature of DOCX editing makes it different and trickier from editing code or HTML. With simple files, the text representation is mostly the thing itself and changes are local. If we take Markdown/simple HTML for example, the structure is at least familiar and styles are local. This problem gets much worse for vertical AI agents.

Companies like Harvey have already run into it

Companies like Harvey have already run into it. When they rebuilt their document editing system, the diagnosis they landed on was that they'd been asking one agent to be both a legal assistant and a Word state machine at the same time.

An agent like Harvey's already has a hard

An agent like Harvey's already has a hard job. It has to read the counterparty's redlines, apply the firm's playbook, check that a defined term means the same thing in clause 3 as it does in clause 27, and catch that the indemnification cap it just changed contradicts the liability section two pages up. That's the work. Splitting runs and chasing numbering references is not, but it competes for the same context window.

Going back to the solutions above, each of

Going back to the solutions above, each of them have different trade-offs, but they all land in the same place: the agent spends its context budget on Word mechanics instead of the actual task. Low-level libraries (python-docx, aspose, Open XML SDK)

News

Launching Vespper DOCX MCP: 3× faster, 2× cheaper, more accurate

Word documents are everywhere.

@spots
Source: Hacker News
See more like this