Skip to content

Read

Site modules

Save what your agent learns about a site, as data, and reuse it on every read.

When your agent works out how a site really works, it can save that as a site module instead of writing a throwaway parser. Typical lessons:

  • the search URL only takes q;
  • results sit as JSON inside the page;
  • listing URLs live in JSON-LD and need pairing with the cards.

Every later read of that site comes back shaped: items with the fields you named, the next page, and checks that tell you when the site changes.

A module is data, never code, and lives in your own state folder. Frankensurf ships none; site knowledge stays with whoever learned it.

Pass auto_items (MCP: items=True) on any read. If the page carries a list of items in its own data, the result gets items, found the same way as below and with no extra read, plus auto_module: the module that found them. Save it (site_modules put) and every later read of that kind of page is shaped without asking.

page = await web.read(url, policy_overrides={"auto_items": True})
page["items"][:3]; page["receipt"]["auto_items"] # {"found", "count", "source"}

Most listing pages carry their rows as data for their own frontend: JSON-LD, JSON in a script tag (__NEXT_DATA__ and friends), or JSON the page fetches while it renders. When none of that holds the rows, they are usually repeated HTML cards. module discover reads a page, looks in all four places, and drafts a module from the best list of items it finds (name, link, price, image). Each draft is validated and run against that same page, so you see the item count and sample rows before you save anything.

found = await web.discover_module("https://www.example.com/search?q=lamp")
best = found["drafts"][0] # module, count, fields, sample
await web.discover_module("https://www.example.com/search?q=lamp", save=True)

After saving, every read of a matching URL (another query, the same path) comes back with items. Check the sample first: a draft can pick the wrong list on a page with several, and nothing here knows a site. On 16 listing pages it had never seen, it drafted the right list on 8, the wrong one on 1, and nothing on 7 (mostly pages it could not read). Keep the drafts that are right, and edit or write the rest by hand.

Start from a home page instead, with a query, when you don’t know the site’s search URL. Frankensurf finds the search itself: schema.org SearchAction first, then the page’s search form, then, for a search box run by script, it types the query once into the visible box in a local browser, presses Enter and keeps the URL it lands on. Nothing else is typed or clicked.

Terminal window
frankensurf module discover https://www.example.com/ --query "desk lamp" --save

On 14 home pages it had never seen, it found the search on 10 (three of the others refused the read) and drafted the right results on 6.

When the page was a search, the draft also gets a search template built from its URL: the parameter that held the search words (?q=lamp), or the path segment after /q/, /search/ or /tag/. Everything else in the URL stays as it was. So one saved search page becomes any search:

page = await web.read_template("auto-example-com", "search", {"query": "desk lamp"})
{
"id": "example.search", "version": "1",
"match": { "origin": "https://www.example.com", "path_pattern": "/search" },
"templates": { "search": { "url": "https://www.example.com/search?q={query}",
"params": { "query": "required" } } },
"sources": { "hits": { "kind": "embedded_json", "decode": "next_flight", "marker": "\"hits\":" } },
"items": { "from": "hits", "fields": {
"title": "ObjectTitle",
"url": { "path": "ITEM_URL", "type": "url" },
"price": { "path": "ObjectPrice", "type": "number" } } },
"assertions": { "schema": "frankensurf.workload-assertions/v1", "checks": [
{ "path": "items", "operator": "minimum_count", "value": 5 },
{ "path": "matching_query", "operator": "minimum_count", "value": 3 } ] },
"notes": "Search takes q; searchTerm is silently ignored."
}
Part What it does
match origin and a path_pattern (a regular expression over the URL path), like a route recipe. Optional query_keys.
templates URLs with {params}, each "required" or {default, encoding: "query" | "path"}. They always stay on the module’s origin.
readiness selector to wait for (bounded), settle_ms.
sources Where rows live: jsonld (type, path), embedded_json (selector, marker, path, optional decode: "next_flight" for Next.js pages), captured_json (JSON the page fetched: url_contains, path), or html (item_selector and fields, each a selector and optional attribute or list of attributes tried in order).
items A field map over one source: name: "path" or {path, type} with type auto, text, number or url. join fills blanks from a second source, by position or by {left, right} key. limit, default 200.
pagination {param, start, step, max_pages} or {next_selector}. The read returns next_url.
invalid Markers of the site’s error or empty-feed page: final_url_contains, title, text. Such a page fails as NOT_FOUND instead of climbing to stronger tools.
assertions A frankensurf.workload-assertions/v1 set over {items, count, first, matching_query}. Operators are nonempty, equals and minimum_count.
policy_defaults Operational settings, such as scroll_screens or settle_ms.

What a module can never do: name an identity, pick or allow providers, allow paid tools, or reach another origin. Any setting you pass on the call wins over the module’s.

  • Saved modules: an enabled one that matches the URL shapes the read by default.
  • One read: name a module with module, run an unsaved one with module_override, or turn modules off with module=False (none on the CLI and MCP).
  • What comes back: items and, with pagination, next_url. The receipt gains module: its id, version, sha256, how it was chosen, the item count and the assertion results (status: passed, failed or invalid). Raw page data is unchanged.
  • Completeness: a page whose module assertions pass counts as complete, so a site that keeps its results in JSON isn’t escalated for having few links.
web.site_modules.put(module) # save or replace (bump version)
page = await web.read_template("example.search", "search", {"query": "desk lamp"})
for item in page["items"]:
print(item["title"], item["price"], item["url"])
print(page["receipt"]["module"]["status"]) # passed / failed / invalid
await web.read(url, module=False) # without modules
web.site_modules.enable("example.search", False) # or disable it

When a module’s assertions start failing, receipt.module.status says failed (or invalid), receipt.module.repair says what to do, and the raw page is still there. Your agent can either:

  • Save a fix directly. Put the module again under a new version. Fine for your own modules.

  • Propose a fix through repair. Use this when changes should be reviewed. The agent drafts the next version (same id and origin, new version) and proposes it with the failing read’s trace_id. Frankensurf then:

    1. checks it with the old version’s assertions, so a fix can’t pass by weakening the checks;
    2. runs it over the page the failing read kept;
    3. runs it over a fresh, independent read;
    4. records the proposal and its validation. Nothing changes yet.

    The owner promotes it with repair-promote. Promotion re-checks the kept page, refuses if the module changed in the meantime, saves the new version and records it in the repair history. repair-disable rolls back to the version it replaced.

proposal = await web.propose_module_repair(failing_trace_id, next_version)
# proposal["status"] == "proposal_ready" once the kept page and a fresh read pass
web.promote_repair(proposal["proposal"]["id"], proposal["proposal"]["sha256"]) # owner
web.disable_repair(proposal["proposal"]["id"], "rolled back after review") # undo