Extract HTML, text, and links
Extract HTML, text, and links.
Already have a browser page?
After navigating with Playwright, extract only what the next step needs:
const html = await page.content();
const text = await page.locator('body').innerText();
const links = await page.locator('a[href]').evaluateAll((anchors) =>
anchors.map((a) => ({
text: a.textContent?.trim() ?? '',
url: new URL(a.getAttribute('href') ?? '', document.baseURI).href,
}))
);
console.log(JSON.stringify({ html, text, links }));page.content() serializes the current DOM, not the original HTTP response bytes. innerText() depends on rendered layout. new URL(..., document.baseURI) resolves relative links against the document base. Limit output size and avoid capturing private account data unnecessarily. This example targets Chromium.
No interactive session needed?
Use Web Fetch extract for its documented extraction result, or Dump DOM for HTML. These are separate requests with their own options and authentication; they do not automatically reuse a browser session's cookies or proxy.
Choose an engine and timeout using the endpoint reference. An HTTP-only fetch will not execute client JavaScript. DOM snapshots may contain scripts/styles unless you request filtering. A captured dom_id can be reused only according to the extraction endpoint's documented lifetime and service availability; it is not a persistent browser Context.
For agents, the CLI exposes bounded snapshots and content extraction. Do not label arbitrary HTML as Markdown or claim a full browser accessibility tree from a DOM-based snapshot.
