Browser-Process Accessibility Tree Cache

Decision: A one-time architectural or governance choice whose consequences still govern current work.

Chromium keeps a complete accessibility-tree cache in the browser process so synchronous operating-system queries can be answered without waiting on a sandboxed renderer.

A screen reader doesn’t ask a page one question. It may ask thousands while building a virtual view of a document: which node is a heading, what text belongs to it, where it sits, and which control it labels. Chromium answers those synchronous questions from a browser-process copy of the page’s semantic structure. The renderer sends updates ahead of time rather than joining each query while an assistive-technology client waits.

Decision Statement

The Chromium project decided to maintain a complete, internally consistent accessibility tree in the browser process and update it asynchronously from renderer-generated tree changes, instead of issuing blocking renderer calls for native accessibility queries.

Context

Blink owns the document object model, style, and layout that give a page semantic meaning. That work runs in a sandboxed renderer process. Native accessibility APIs on Windows, macOS, Linux, ChromeOS, and Android live on the privileged side of the browser-renderer boundary, and many of those APIs expect synchronous answers.

An accessibility tree is a purpose-built semantic view of a page. It exposes roles, names, states, relationships, and bounds to assistive technology without exposing the whole document model. Blink derives a per-frame tree through AXObjectCache. After layout reaches a clean lifecycle state, AXTreeSerializer emits batched AXTreeUpdate changes. The browser applies those changes to its own AXTree and presents native objects through BrowserAccessibilityManager and AXPlatformNode.

Site Isolation makes the arrangement more than a two-copy cache. Each frame has its own tree in its renderer. The browser connects parent and child frame trees with AXTreeID values and presents the result as one virtual tree. Those identifiers are capability-bearing: a compromised renderer must not discover or forge an embedder token for an unrelated tree.

Alternatives Considered

AlternativeWhat it would doWhy Chromium did not choose it
Per-query proxy IPCKeep little accessibility state in the browser and make a blocking renderer call for each native query.Screen readers may issue thousands of synchronous calls while constructing a virtual buffer. The project’s accessibility history records page loads approaching ten seconds, plus deadlock risk whenever layout or the renderer couldn’t answer.
Direct native exposure from each rendererLet a renderer host the operating-system accessibility objects for its own documents.A renderer lacks the operating-system authority needed for native integration. The design would also push privileged, synchronous API work into the process that parses untrusted web content and would fragment cross-frame trees under Site Isolation.
Complete browser-process cache (chosen)Push batched semantic updates to a complete browser copy and answer native queries there.This spends memory and accepts short-lived staleness, but native calls see a consistent snapshot without blocking on a renderer.

The rejected proxy was attractive because it avoided duplicating the tree. Its latency scaled with the number of questions rather than the amount of semantic change, though, and a renderer waiting on layout could block the browser while the browser waited for an answer. The chosen design moves that work off the query path.

Rationale

The decision gives the synchronous API an asynchronous data source. Renderer work can be batched after Blink reaches a safe lifecycle point, while the browser answers native calls from the last complete tree it accepted. A slow or crashed renderer may leave the copy briefly stale, but it can’t stall a screen reader’s browser-side query loop.

The data flow crosses the privilege boundary once per update batch, not once per native query:

flowchart LR
  DOM[Document and layout] --> RT[Renderer semantic tree]
  RT --> U[Atomic tree update]
  U --> BT[Browser tree cache]
  BT --> API[Native accessibility API]
  API --> AT[Assistive technology]

Diagram: Flowchart. The diagram contains these connections:

• Document and layout leads to Renderer semantic tree.

• Renderer semantic tree leads to Atomic tree update.

• Atomic tree update leads to Browser tree cache.

• Browser tree cache leads to Native accessibility API.

• Native accessibility API leads to Assistive technology.

Atomic application matters as much as caching. AXTreeUpdate is stateful: an update only makes sense against the tree version the browser already holds. The browser applies the update through the canonical AXTree deserializer, rejects an update that doesn’t fit, and can reset the stream when the two sides diverge. A native client sees the old complete tree or the new complete tree, never a half-applied mutation.

The cache also keeps operating-system objects out of sandboxed renderers. Platform-specific wrappers and synchronous callbacks terminate in browser-side code. Renderer code supplies semantic data, not native authority.

Ongoing Consequences

The browser pays for two representations of the page’s semantics. Renderer and browser memory rise with document size, and the browser copy can lag DOM or layout changes until the next clean update. AXMode, targeted accessibility, and progressive activation limit the work when no consumer needs the full tree, but an active assistive-technology client can require substantially more semantic state.

Every update is untrusted. Mojo can prove that an AXTreeUpdate has the declared wire shape; it can’t make node IDs, text, roles, bounds, relationships, or child counts truthful. Browser-side code must use the canonical tree-update path, validate sensitive identifiers, and keep arbitrary processing of web-controlled accessibility data out of privileged C++ where possible. Chromium’s security guidance routes new complex processing to a sandboxed process or a memory-safe language.

Stateful replication is a bounded exception to the default Stateless IPC Interface rule. The renderer remembers which nodes its counterpart has seen and sends deltas. That state improves efficiency, but it must never confer privilege. Browser-owned origin checks, frame ownership, tree identifiers, and access to sensitive content remain authoritative on every use.

Site Isolation adds composition work. The browser must join frame trees without exposing one renderer’s protected tree identifier to another renderer. Actions travel in the opposite direction from updates: the browser routes a focus, scroll, or default-action request to the renderer that owns the target node, and the renderer performs it asynchronously. Hit testing may require a browser-side approximation followed by a renderer-side correction because layout and transforms remain renderer knowledge.

Downstream Chromium-based products inherit both sides of the trade. Removing the cache breaks native accessibility clients that assume synchronous responses. Extending it with product-specific attributes adds browser memory and expands the surface that accepts attacker-controlled content. A fork’s accessibility audit therefore needs memory measurements, update-failure telemetry, and security review together.

Reversal Conditions

The decision remains necessary while native accessibility APIs are synchronous and sandboxed renderers can’t call them directly. It could be revisited if every supported platform adopted an asynchronous accessibility protocol, or if Chromium moved the complete semantic tree into a separately sandboxed service that could answer platform calls without browser-process risk or query-path IPC.

A rewrite would also become plausible if a memory-safe representation and cross-process sharing mechanism eliminated most duplicate storage while preserving atomic snapshots. None of those conditions holds across Chromium’s supported platforms. The current security guidance’s preference for sandboxed or memory-safe processing narrows where new logic belongs; it doesn’t remove the browser cache that existing native APIs depend on.

Notes for Agent Context

Keep native accessibility queries on the browser-side BrowserAccessibilityManager and answer them from its cached AXTree; don’t add synchronous per-query IPC to a renderer. Apply renderer updates only through the canonical AXTree deserialization path, and treat every string, ID, role, relation, offset, and bound as web-controlled after Mojo validation. Preserve browser-owned AXTreeID and frame-ownership checks when joining Site Isolation trees, and never expose an embedder token to an unrelated renderer. Route new complex processing of accessibility data to a sandboxed process or a memory-safe language instead of adding ad hoc privileged C++ over raw AX mojom structures.

Bounded by: Memory Pressure Response — Keeping complete trees in two processes trades memory for bounded query latency and freedom from synchronous renderer calls.

Complements: Rendering Pipeline — The rendering pipeline produces pixels while the accessibility pipeline produces the semantic structure consumed by assistive technology.

Contrasts with: Stateless IPC Interface — Accessibility updates are stateful replication, but renderer-supplied state never becomes authority for a privileged decision.

Depends on: Multi-Process Architecture — The browser-side cache exists because Chromium separates renderer-owned document semantics from browser-owned operating-system integration.

Enforces: Browser-Renderer Privilege Split — Renderers send semantic updates across the privilege boundary instead of exposing native accessibility objects themselves.

Enforces: Untrusted Renderer Axiom — The browser treats every accessibility update as web-controlled data even after Mojo has validated its structure.

Refined by: Site Isolation — Site Isolation gives each frame its own accessibility tree, which the browser joins into one virtual tree with protected tree identifiers.

Sources

Chromium’s How Chrome Accessibility Works, Part 2 records the failed per-query proxy, the full-cache decision, atomic updates, renderer lifecycle gating, and the stateful AXTreeSerializer protocol. The project’s Accessibility Overview maps renderer semantics to the native accessibility APIs on supported platforms.

How Chrome Accessibility Works, Part 3 describes platform abstraction, cross-frame tree composition, asynchronous actions, and hit testing. Chromium’s Security Guidelines for the Accessibility Tree defines accessibility data as web-controlled even after structural Mojo validation and bounds browser-process use. The W3C’s WAI-ARIA 1.2 Accessibility Tree model supplies the standards vocabulary for roles, states, properties, inclusion, and exclusion.

Technical Drill-Down

• docs/accessibility/browser/how_a11y_works_2.md (pinned bf92b3e) — the cache decision, serializer protocol, atomic update path, and rejected blocking proxy.

• docs/accessibility/browser/how_a11y_works_3.md (pinned bf92b3e) — platform wrappers, cross-frame composition, actions, hit testing, and tree identifiers.

• docs/security/ax-tree-security-guidelines.md (pinned bf92b3e) — the trust rules for structurally valid but attacker-controlled accessibility data.

• content/browser/accessibility/ (pinned bf92b3e) — browser-process cache ownership, platform managers, event routing, and diagnostic surfaces.

• ui/accessibility/ax_tree.h (pinned bf92b3e) — the canonical browser-side tree and update application API.