> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mcp-b.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# @mcp-b/smart-dom-reader reference

> Token-efficient DOM extraction for AI agents, with ranked selectors and progressive extraction.

`@mcp-b/smart-dom-reader` extracts interactive controls, semantic content, page structure,
and ranked selectors from browser DOM trees. The package has no runtime dependencies.

```text title="Package metadata" theme={null}
npm: @mcp-b/smart-dom-reader
license: MIT
dependencies: none
```

## Installation

```bash theme={null}
npm install @mcp-b/smart-dom-reader
```

## Minimal example

```typescript title="Extract controls and the main region" theme={null}
import { ProgressiveExtractor, SmartDOMReader } from '@mcp-b/smart-dom-reader';

const interactive = SmartDOMReader.extractInteractive(document);

const structure = ProgressiveExtractor.extractStructure(document);
const mainSelector = structure.summary.mainContentSelector;
const region = mainSelector
  ? ProgressiveExtractor.extractRegion(mainSelector, document, { mode: 'interactive' })
  : null;

console.log({ interactive, region });
```

## Entry points

| Entry point | Exports |
| - | - |
| `@mcp-b/smart-dom-reader` | Extraction classes, formatting utilities, and TypeScript types |
| `@mcp-b/smart-dom-reader/bundle-string` | `SMART_DOM_READER_BUNDLE` and `SMART_DOM_READER_VERSION` |

`SmartDOMReader` is both a named export and the package's default export.

## `SmartDOMReader`

`SmartDOMReader` returns one `SmartDOMResult` for a `Document` or `Element`.
Runtime options override constructor options.

```typescript theme={null}
class SmartDOMReader {
  constructor(options?: Partial<ExtractionOptions>);

  extract(
    rootElement?: Document | Element,
    runtimeOptions?: Partial<ExtractionOptions>
  ): SmartDOMResult;

  static extractInteractive(doc: Document, options?: Partial<ExtractionOptions>): SmartDOMResult;

  static extractFull(doc: Document, options?: Partial<ExtractionOptions>): SmartDOMResult;

  static extractFromElement(
    element: Element,
    mode?: ExtractionMode,
    options?: Partial<ExtractionOptions>
  ): SmartDOMResult;
}
```

`extract()` defaults to the current global `document`. `extractInteractive()` forces
`mode: 'interactive'`; `extractFull()` forces `mode: 'full'`.

## `ProgressiveExtractor`

```typescript theme={null}
class ProgressiveExtractor {
  static extractStructure(root: Document | Element): StructuralOverview;

  static extractRegion(
    selector: string,
    doc: Document,
    options?: Partial<ExtractionOptions>
  ): SmartDOMResult | null;

  static extractContent(
    selector: string,
    doc: Document,
    options?: ContentExtractionOptions
  ): ExtractedContent | null;
}
```

`extractRegion()` and `extractContent()` return `null` when `selector` has no match.
When `extractStructure()` receives an `Element`, that element is reported as the main region.

## Extraction modes

```typescript theme={null}
type ExtractionMode = 'interactive' | 'full' | 'structure' | 'content';
```

| Value | Contract |
| - | - |
| `interactive` | Returns page state, landmarks, forms, and interactive elements |
| `full` | Adds `semantic` and `metadata` to the interactive result |
| `structure` | Compatibility label; direct `SmartDOMReader` output remains interactive-shaped |
| `content` | Compatibility label; direct `SmartDOMReader` output remains interactive-shaped |

`SmartDOMReader` adds semantic output only for `full`. Use `ProgressiveExtractor.extractStructure()`
and `ProgressiveExtractor.extractContent()` for the specialized structure and content result types.

## `ExtractionOptions`

The public interface requires `mode`; every constructor and method accepts
`Partial<ExtractionOptions>` and supplies the defaults below.

| Field | Type | Default | Contract |
| - | - | - | - |
| `mode` | `ExtractionMode` | `interactive` | Result mode |
| `maxDepth` | `number` | `5` | Maximum traversal depth |
| `includeHidden` | `boolean` | `false` | Includes elements that fail visibility checks |
| `includeShadowDOM` | `boolean` | `true` | Recursively queries open shadow roots |
| `includeIframes` | `boolean` | `false` | Accepted compatibility field; the reader does not traverse frames itself |
| `viewportOnly` | `boolean` | `false` | Includes only elements intersecting the viewport |
| `mainContentOnly` | `boolean` | `false` | Uses detected main content when the root is a `Document` |
| `customSelectors` | `string[]` | `[]` | Adds matching elements to `interactive.clickable` |
| `attributeTruncateLength` | `number` | `100` | Maximum selected-attribute value length, including recognized test IDs |
| `dataAttributeTruncateLength` | `number` | `50` | Maximum value length for additional `data-*` attributes |
| `textTruncateLength` | `number` | unlimited | Maximum extracted text length per element |
| `filter` | `FilterOptions` | none | Applies element-level filters |

### `FilterOptions`

| Field | Type | Match rule |
| - | - | - |
| `includeSelectors` | `string[]` | Any selector |
| `excludeSelectors` | `string[]` | Excludes on any selector |
| `textContains` | `string[]` | Any case-insensitive substring |
| `textMatches` | `RegExp[]` | Any regular expression |
| `hasAttributes` | `string[]` | Every named attribute |
| `attributeValues` | `Record<string, string \| RegExp>` | Every attribute-value condition |
| `tags` | `string[]` | Tag name appears in the list |
| `interactionTypes` | `Array<keyof ElementInteraction>` | Any requested interaction flag |
| `withinSelectors` | `string[]` | Element is within any matching ancestor |
| `nearText` | `string` | Parent text contains the value, case-insensitively |

### `ContentExtractionOptions`

| Field | Type | Default | Contract |
| - | - | - | - |
| `includeHeadings` | `boolean` | `true` | Includes H1-H6 headings |
| `includeLists` | `boolean` | `true` | Includes ordered and unordered lists |
| `includeTables` | `boolean` | `true` | Includes table headers and body rows |
| `includeMedia` | `boolean` | `true` | Includes image, video, and audio references |
| `preserveFormatting` | `boolean` | — | Reserved public field; the current extractor does not consult it |
| `maxTextLength` | `number` | unlimited | Maximum length of headings, paragraphs, and list items |

### `MarkdownFormatOptions`

| Field | Type | Contract |
| - | - | - |
| `detail` | `'summary' \| 'region' \| 'deep'` | Reserved field; current formatter methods do not inspect it |
| `maxTextLength` | `number` | Maximum rendered text length per item |
| `maxElements` | `number` | Maximum rendered items in each repeated group |

## Result types

### `SmartDOMResult`

| Field | Type | Contract |
| - | - | - |
| `mode` | `ExtractionMode` | Effective result mode |
| `timestamp` | `number` | Extraction start time in Unix milliseconds |
| `page` | `PageState` | URL, title, loading, error, modal, and optional focus state |
| `landmarks` | `PageLandmarks` | Selectors grouped by detected landmark |
| `interactive` | object | Buttons, links, inputs, forms, and other clickable elements |
| `semantic` | object | Headings, images, tables, lists, and articles; `full` mode only |
| `metadata` | object | Element counts, main-content selector, and language; `full` only |

### Element types

| Type | Contract |
| - | - |
| `ExtractedElement` | Tag, text, selector, attributes, context, interaction flags, and children |
| `ElementSelector` | Best CSS selector, XPath, optional hints, and ranked candidates |
| `ElementSelectorCandidate` | `type`, `value`, and numeric `score` |
| `ElementContext` | Nearest form, section, main, nav, and parent chain |
| `ElementInteraction` | Compact optional flags; booleans are emitted only when `true` |
| `FormInfo` | Form selector, action, method, inputs, and buttons |
| `PageState` | Current page URL, title, diagnostic flags, and focused selector |
| `PageLandmarks` | Navigation, main, form, header, and footer selectors; `articles` and `sections` share detected region selectors |

### Progressive result types

| Type | Contract |
| - | - |
| `StructuralOverview` | Regions, form summaries, page summary, and optional extraction suggestions |
| `RegionInfo` | Region selector, label, role, feature flags, counts, and text preview |
| `ExtractedContent` | Headings, paragraphs, lists, tables, media, word count, and interaction flag |

## Selector generation

```typescript theme={null}
class SelectorGenerator {
  static generateSelectors(element: Element): ElementSelector;
  static getContextPath(element: Element): string[];
}
```

`generateSelectors()` sorts `candidates` from highest to lowest score and selects the
highest-ranked CSS candidate for `ElementSelector.css`.

| Candidate | Score | Notes |
| - | - | - |
| `id` | `100` | Emitted only for a unique ID |
| `data-testid` | `90`, plus `5` when unique | Recognizes common test-ID attribute names |
| `role-aria` | `85`, plus `5` when unique | Requires both `role` and `aria-label` |
| `name` | `78`, plus `5` when unique | Uses the `name` attribute |
| `class-path` | `max(0, 70 + class bonus - structural penalties)` | Class bonus is `8`; each `:nth-child` is `-10` |
| `xpath` | `40` | Fallback XPath |
| `text` | `30` | Hint for short button, link, or label text |

The `ElementSelectorCandidate.type` union also contains `css-path`; the current generator
emits generated CSS paths as `class-path`. Text candidates use the non-standard `:contains()`
form as a hint and are never selected for `ElementSelector.css`.

## Supporting classes

```typescript theme={null}
class ContentDetection {
  static findMainContent(doc: Document): Element;
  static calculateContentScore(element: Element): number;
  static isNavigation(element: Element): boolean;
  static isSupplementary(element: Element): boolean;
  static detectLandmarks(doc: Document): {
    navigation: Element[];
    main: Element[];
    complementary: Element[];
    contentinfo: Element[];
    banner: Element[];
    search: Element[];
    form: Element[];
    region: Element[];
  };
}

class MarkdownFormatter {
  static structure(
    overview: StructuralOverview,
    options?: MarkdownFormatOptions,
    meta?: { title?: string; url?: string }
  ): string;

  static region(
    result: SmartDOMResult,
    options?: MarkdownFormatOptions,
    meta?: { title?: string; url?: string }
  ): string;

  static content(
    content: ExtractedContent,
    options?: MarkdownFormatOptions,
    meta?: { title?: string; url?: string }
  ): string;
}
```

The formatter methods return XML-wrapped Markdown strings.

## Bundle string API

The `bundle-string` entry point exposes an injectable IIFE string and its bundle-format
version marker. After evaluating `SMART_DOM_READER_BUNDLE`, call:

```typescript theme={null}
SmartDOMReaderBundle.executeExtraction(
  ...request: { [M in ExtractionMethod]: [method: M, args: ExtractionArgs[M]] }[ExtractionMethod]
): ExtractionResult;
```

The `method` literal selects the argument type. `ExtractionResult` is the formatted `string` or
`{ error: string }`.

| `ExtractionMethod` | Method-specific arguments |
| - | - |
| `extractStructure` | `selector?` |
| `extractRegion` | `selector`, `mode?`, `options?` |
| `extractContent` | `selector`, `options?` |
| `extractInteractive` | `selector?`, `options?` |
| `extractFull` | `selector?`, `options?` |

Every argument type also accepts `frameSelector?` and `formatOptions?`. `frameSelector`
must identify a same-origin iframe with an accessible `contentDocument`.

The root package exports `ExtractionMethod`, `ExtractionArgs`, `ExtractionResult`, and the
five method-specific argument types: `ExtractStructureArgs`, `ExtractRegionArgs`,
`ExtractContentArgs`, `ExtractInteractiveArgs`, and `ExtractFullArgs`.

## MCP server

The separate `@mcp-b/smart-dom-reader-server` package wraps the browser library with
`@modelcontextprotocol/server` v2 and returns XML-wrapped Markdown.

| Tool | Parameters |
| - | - |
| `browser_connect` | `headless?`, `executablePath?` |
| `browser_navigate` | `url` |
| `dom_extract_structure` | `selector?`, `detail?`, `maxTextLength?`, `maxElements?` |
| `dom_extract_region` | `selector`, `options?` |
| `dom_extract_content` | `selector`, `options?` |
| `dom_extract_interactive` | `selector?`, `options?` |
| `browser_screenshot` | `path?`, `fullPage?` |
| `browser_close` | none |

## Related

* [Chrome DevTools MCP](https://github.com/ChromeDevTools/chrome-devtools-mcp)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.