Skip to content

Read

Page data

Get the structured data a page already carries, including the JSON its own frontend fetched.

Most pages carry their data in machine-readable form before any text extraction. Frankensurf hands it over as-is, so your agent can read prices, dates or listings without scraping the layout.

  • structured.jsonld: every JSON-LD block the page publishes.
  • structured.embedded_json: JSON embedded for the page’s own frontend, such as Next.js’s __NEXT_DATA__.
  • captured_json: on rendered reads, the JSON responses the page fetched from its own API.
page = await web.read("https://shop.example.com/p/2231", policy_overrides={
"render": True,
"capture_json_responses": True,
})
offers = [item["data"] for item in page.get("captured_json", {}).get("items", [])]
Response (trimmed)
{
"structured": {
"jsonld": [{ "@type": "Product", "name": "…", "offers": { "price": "…" } }],
"embedded_json": []
},
"captured_json": {
"items": [{
"url": "https://shop.example.com/api/product/2231",
"http_status": 200,
"content_type": "application/json",
"format": "json",
"data": { "title": "Widget", "price": 42 }
}],
"skipped": 0
}
}

When the URL itself returns JSON, structured is the parsed body. Newline-delimited JSON becomes structured.records.

Option Default Description
capture_json_responses false Capture the page’s own JSON responses (rendered and signed-in reads).
capture_json_max_items 20 Most responses kept.
capture_json_max_bytes 2 MB Largest response kept.