Skip to content

HTTP Archive Documentation

Guides, schema references, and query recipes for analyzing how the web is built.

Start Querying BigQuery

New to HTTP Archive? Learn how to access the public httparchive dataset on Google BigQuery and run your first query in the Getting Started guide.

Minimizing Query Costs

HTTP Archive crawls petabytes of web data. Learn partitioning, clustering, and cost-control best practices in the Minimizing Costs guide.

Guided Tour

Walk through real-world analysis examples, query patterns, and summary tables in our Guided Tour.

Release Cycle & Methodology

Learn how millions of URLs are crawled each month, how pipelines process data, and when new datasets are published in the Release Cycle guide.

Pages & Requests Tables

Detailed column specifications, clustering keys, and JSON payload definitions for the core pages and requests tables.

Custom Metrics

Explore custom metrics extracted during crawls, from ad tracking to privacy and other runtime features.

Functions & Structs

Inspect schema structs like technologies and features, plus user-defined functions like Capo head analysis.

Raw Payloads & Blobs

Understand how Lighthouse audits, page summaries, and request payloads are formatted and stored.

The HTTP Archive and its documentation are open-source. Found a typo, have a query recipe to share, or want to improve custom metrics? Contribute on GitHub.