PDF Importer
We pay Datalab once, our server does the conversion, and every Seed agent gets high quality PDF imports out of the box without touching the key

Problem

Importing PDFs already works well. The Datalab importer gives much better output than what the agent can do alone in the sandbox. The problem is that to use it you need a Datalab account, an API key, and put that key in a file in the agent memory. Almost nobody will do that.

What I want: the user creates an agent, drops some PDFs in the memory, says "import these", and gets good Seed documents. No other account, no key, nothing to configure. We pay one Datalab subscription for everybody.

The catch is the key. We can't give it to the agents: the model reads whatever the agent has, memory can be shared with collaborators, and for most users the agent runs in their own laptop. So the Datalab call has to happen in a server we run, and the key never leaves that server. And still every agent, wherever it runs, has to be able to import.

The idea

The server converts, the agent reviews and publishes.

The user drops paper.pdf in memory and asks the agent to import it. The agent calls a new tool, convert. The server takes the PDF from the memory folder, sends it to Datalab with our key, waits, and writes the markdown and images back in memory next to the PDF. The agent gets one line back, something like "converted, 18 pages, 6 images, see datalab-imports/paper/seed.md". Then it reads the markdown, fixes what needs fixing (bibliography, sections, links) and publishes it like today.

The agent never talks to Datalab, never sees the key, never touches the bytes.

What convert is

Every tool an agent has is a tool document. There are three kinds: built-in tools live in the agents server itself (search, query, attributes, web_search and execute are the ones we have), authored tools are snippets the agent or the user wrote and run in the sandbox, and MCP tools come from a remote MCP server the account connected.

convert is a new built-in, same family as web_search. The model calls it with call like any other tool and the server runs it. There is no code in memory, the model only sees the tool document: name, description and input schema.

The inputs are the ones the Datalab importer reference already has: pdfs, output_dir, page_range, max_pages, overwrite, the structured extraction fields. Same output folder too (raw.md, document.md, seed.md, manifest.json, assets/). The only thing that goes away is api_key_file. So agents that already read that reference keep working, we just edit the reference.

Why this works with what we have

This is less new than it sounds.

The agents server is the agent runtime. The model only returns text and tool calls, the server executes everything: read, write, search, execute, all of it. convert is one more case. That is why the key is safe, the model only sees the result, same as it sees search results but never the search backend.

Every request to an agents server is signed. The app signs each action with the account key through the daemon and the server checks the signature before doing anything. So the server always knows which account is asking, and that is all we need to count pages.

Agents find their tools alone. The system prompt has a short index with one line per tool document, name and summary. A new built-in shows up there for every agent, and the agent can read ~/tools/convert when it wants details. If it calls with a wrong input it gets the tool document back. The Agent Guide only has to say "for PDFs, call convert".

The server already keeps encrypted secrets per account (model keys, MCP headers). A user with their own Datalab key saves it there and their conversions don't count against the shared allowance.

And it is the same binary everywhere: our hosted server, the local server the desktop app starts in your laptop, a self hosted Docker. So convert exists everywhere from day one. What is not everywhere is the key.

The modules

The converter. Takes PDF bytes and a key, calls Datalab, polls until done, returns markdown and images. Nothing else in the server knows about Datalab.

The key chooser. If the account has its own Datalab key, use it and count nothing. If not, use our shared key while the account has allowance. If not, answer with a clear message on how to add your own key.

The ledger. A small table with pages used per account per month, plus a global counter for the shared key. Some free pages per account, and a hard ceiling on the shared key so a bad month can't cost more than we decided. Pages are reserved before the call and settled after, so a crash in the middle doesn't make us pay twice.

The tool document, the guide update, and a flag in the server health so the app can show if importing is available.

The remote door. A new signed action in our hosted server, ConvertDocument: send a PDF, get markdown and images back. This is what makes it work for agents that don't live in our server.

Who pays

Keep it simple, no new permissions.

In our hosted server the owner is known, the agent belongs to an account, so the owner's allowance is charged.

When the request comes from another server, it is signed with the agent's own identity key. Every agent has one, created with the agent, and it is a normal Seed account. That account is charged. We don't know who created the agent from our side (the published profile only has a display name), but if the owner published a capability to that agent key, we can read that public blob and charge the owner instead. Just a lookup to know who to bill, never something that gives that key any power in our server.

Yes, someone can create more agents to get more free pages. The global ceiling bounds our cost, good enough to start.

Where it goes

Important thing: a desktop user's agents don't run in our hosted server. The desktop app starts its own agents server in the laptop and the agents live there. So something that only works in our hosted server is useless for the main case.

That is why the remote door is in the first version. convert is in the binary, so an agent created in a laptop has it the moment it exists. The key is not in the binary, so when that agent calls convert, its local server asks our hosted server through ConvertDocument, gets the result, and writes it in the local memory. For the model nothing changes. Same for self hosters.

So the first version is: built-in tool, converter, key chooser, ledger, ConvertDocument, the relay in the local server, health flag and guide update. With that every agent anywhere can import PDFs.

As a side effect, anything that signs a Seed request can call ConvertDocument with its own key: seed-cli, a Claude Code session, whatever.

Structured extraction

Datalab can also extract fields from the PDF into a schema you pass, reusing the conversion so the PDF is not parsed twice. The importer reference already has this (structured_fields, structured_schema, structured_schema_id), so it comes with the same tool.

The interesting part is Hypermedia Schemas. If convert accepts the hm:// URL of a schema page, the extraction fills the attributes of a typed document straight from the PDF and our validation checks them. Import a folder of papers into a "Paper" type with title, authors and year filled, without the model guessing.

Datalab charges this separately and it is more expensive: conversion is $4 per 1,000 pages, extraction goes from $6 (fast) to $15 (balanced) and $20 (accurate). So it is a second line in the ledger, and it comes after basic conversion is out.

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime