pdfmend

Guides / How-to

How to automate PDF editing with AI agents and WebMCP

Updated 2026-08-27 · pdfmend team

Most PDF automation sends a document to a remote service. pdfmend takes a different approach: the editor runs inside the browser, and its automation surfaces call the same local editing logic as the visible controls. A person can use the interface, a Playwright-class agent can call a versioned browser API, and a compatible browser agent can discover WebMCP tools. The document still does not need a pdfmend upload endpoint.

That makes pdfmend agent-ready in a specific, testable sense. It publishes machine-readable discovery files, stable selectors, structured actions, and explicit error codes. It does not mean that every AI assistant can control every feature automatically, or that WebMCP is already available in every browser. This guide explains the working surfaces and their boundaries.

What agent-ready means in pdfmend

An agent-ready site needs more than prose saying that AI is supported. An agent must be able to discover what the product does, understand the inputs to an action, observe whether the editor is ready, and receive a structured result or error. pdfmend provides four complementary layers:

  • Accessible browser UI. Semantic controls, ARIA labels, and stable test identifiers let browser-driving agents use the same interface as a person.
  • Versioned browser API. window.pdfmend.v1 exposes structured editor actions for trusted Playwright, Puppeteer, Stagehand, or similar automation.
  • WebMCP tools. When a compatible browser or bridge provides document.modelContext, pdfmend registers typed, metadata-only tools that a browser agent can discover and invoke.
  • Machine-readable documentation. llms.txt, llms-full.txt, agents.json, and the Agent Skills discovery manifest describe the product, interfaces, privacy contract, and stable automation flow.

These layers are useful independently. A current automation runner can use the window API today without waiting for native WebMCP support, while an experimental WebMCP client can use the safer tool subset.

Choose the right automation surface

Use the visible UI when the workflow depends on a control that is not in the programmatic contract, or when a human should review each step. The automation section in llms.txt documents the ready signals, roles, and stable selectors for opening a document, editing it, and exporting a copy.

Use window.pdfmend.v1 for trusted browser automation that needs to open PDF bytes, inspect page geometry, add text or images, reorganize pages, undo or redo, and receive an exported copy. The API accepts and returns base64 for document operations, so the controlling automation environment must be one you trust. pdfmend does not transmit those values, but caller code controls what happens after it receives them.

Use WebMCP when the browser agent should operate through a deliberately narrow tool contract. The WebMCP facade never accepts document bytes and never returns them in a tool result. It can report capabilities and state, create a blank document, add text, trigger a local download, or save a copy to the browser's local library. This keeps PDF content out of a tool response that may be sent to a cloud model.

Discover the contract before acting

Start with https://pdfmend.app/agents.json. It identifies the editor URL, the versioned window handle, the WebMCP tool names, and the privacy terms in a small JSON document. Read https://pdfmend.app/llms.txt for the concise automation and scripting reference. llms-full.txt adds the public guide corpus, while /.well-known/agent-skills/index.json points to a digest-pinned skill for agents that support Agent Skills discovery.

On the landing page, a compatible WebMCP client can call pdfmend_get_capabilities to receive the same discovery pointers. It can also call pdfmend_get_version to bind evidence to the exact deployed build. On the editor, the window API reports version: 1, and WebMCP adds the editor-specific tools only while that editor surface is mounted.

Treat ready, busy, documentOpen, dirty, and signatureGate as state, not decoration. An agent should wait for readiness, avoid concurrent edits, and handle a signed-source acknowledgement or dirty-session decision instead of clicking through warnings.

Automate a PDF safely, step by step

  1. Open the editor in a browser context controlled by the user or by a trusted local automation runner.
  2. Wait for the editor API to exist and for getState() to report ready. Do not rely on a fixed sleep or on the page looking visually complete.
  3. Choose the interface. Use window.pdfmend.v1 when document bytes must enter or leave the automation context; use WebMCP for its metadata-only action set; use stable UI controls for features outside either facade.
  4. Read state before mutation. Page numbers are one-based, and text/image coordinates use PDF points with the origin at the bottom-left of the unrotated page.
  5. Apply one or more supported actions, then read state again. Handle stable error codes such as SESSION_BUSY, SESSION_DIRTY, SIGNATURE_GATE, or PDF_PAGE_NOT_FOUND rather than guessing from visible text.
  6. Export deliberately. The window API can return base64 to the trusted runner. WebMCP instead triggers a local browser download and returns only the file name and byte length. A local-library save also returns metadata only and requires storage consent.
  7. Verify the output in the same browser workflow before sharing it. Agent access does not turn experimental editing features into certified redaction or remove the need to inspect important documents.

WebMCP tools available today

The landing page registers pdfmend_get_capabilities and pdfmend_get_version. The editor registers pdfmend_get_state, pdfmend_open_blank, pdfmend_add_text, pdfmend_export_download, and pdfmend_save_to_library. Every result is deliberately small and contains metadata rather than PDF bytes.

WebMCP is an evolving W3C Community Group draft, not a universally deployed browser API. pdfmend feature-detects the current document.modelContext surface and keeps a compatibility path for the earlier navigator surface. It does not ship a WebMCP polyfill. In an ordinary browser without a native API or bridge, the WebMCP tools simply do not register; the site and the versioned window API continue to work normally.

Privacy and capability boundaries

The core promise remains the same as in the local editing guide: pdfmend has no server endpoint that receives editor document contents. UI actions and programmatic actions run inside the browser. Content-free, cookieless product analytics may still be sent, but PDF bytes and text are not analytics fields.

The two automation facades intentionally have different trust models. The window API is powerful enough to move base64 between the page and a trusted runner. WebMCP is safer for a browser agent backed by a remote model because its tool contract excludes document bytes in both directions. Neither facade is a general remote PDF-processing API.

The current APIs also do not expose every toolbar feature. There is no WebMCP tool for importing an existing PDF, images, arbitrary page operations, OCR, or bulk text replacement. Large base64 documents create substantial memory overhead, and unusual PDF page boxes are best-effort in version 1. Use the UI or stable selectors when a supported workflow is not yet a structured verb, and keep programmatic documents modest in size.

Try it now — the editor runs entirely in your browser. No upload, no account, no watermark.

Open the free editor

FAQ

Is pdfmend an MCP server?

Not in the conventional remote-server sense. pdfmend implements the WebMCP draft inside its web pages and publishes machine-readable agent discovery. The tools call browser-local product logic; there is no server that receives or edits the PDF.

Can an AI agent open my existing PDF through WebMCP?

No. The metadata-only WebMCP contract intentionally has no document-byte input. A trusted Playwright-class runner can use window.pdfmend.v1, or an agent can drive the browser's visible file-opening flow with the user's authorization.

Does agent automation upload the document?

pdfmend's automation code does not upload it. The WebMCP facade excludes PDF bytes completely. The window API passes base64 to the controlling runner, so you must trust that runner and its surrounding code not to transmit the data.

Which interface is stable?

The browser API is explicitly versioned as window.pdfmend.v1, and its error codes and semantics are tested as a contract. WebMCP follows an evolving draft and is feature-detected, so its browser integration may change as the standard matures.

Related guides