Public research release · v1.0.0

50,000 domains, categorized for research

A documented, reproducible sample from the Web Filtering Database: 1,000 popularity-ranked domains for recognizable inspection and 49,000 domains sampled across category and popularity strata.

RELEASE COMPOSITION
1,000popularity-ranked domains
49,000stratified domains
50,000unique domains
59taxonomy categories
6sampling bands
20.1Msource records analyzed
What this release is

A useful research sample, with its selection method left visible

The release is large enough for analysis and software testing, but small and static enough to remain clearly separate from the complete production database.

Inspectable

The first segment contains domains people are likely to recognize. Researchers can examine concrete classifications instead of starting with thousands of unfamiliar long-tail names.

Broadly distributed

The second segment reaches across six source-rank bands and primary categories, providing substantially more variety than a list made only from the most popular sites.

Reproducible

The package includes the sampling program, fixed random seed, allocation table, machine-readable analysis, checksums, taxonomy files and a complete data dictionary.

Why two components?

Recognition and representation are different problems

If a sample contains only famous domains, it is easy to understand but biased toward large platforms and heavily visited content. Those domains are also more likely to support several functions from the same destination. If a sample is drawn uniformly from millions of records, it reaches the long tail but can become difficult to inspect and may provide too few examples of small categories.

This release does not hide that tradeoff. The first 1,000 unique domains follow the popularity order of the source extract. The remaining 49,000 are selected separately across popularity bands and primary-category strata. Every row states which component and band it came from.

Important interpretation ruleUse the top 1,000 for recognizable examples and comparison. Use the stratified 49,000 for broader testing. Do not treat the combined file as a simple random estimate of the entire internet.

Sample architecture

The commercial database currently covers up to 120 million domains. This release was generated from a 20,128,938-record source snapshot.

Sampling method

49,000 domains distributed across six popularity bands

Each band is subdivided by the first ordered category label. Smaller available categories receive protected representation before the remaining capacity is allocated proportionally.

high visibility
4,000
6,000
9,000
10,000
10,000
long tail
10,000

What “stratified” means here

Within each rank band, domains are grouped by `primary_category`. The procedure first protects up to 25 observations for each available category stratum. It then distributes the remaining places in proportion to the sizes of the remaining strata and selects rows using reservoir sampling.

The fixed seed is 20260922. Duplicate domains are removed: the earliest selected occurrence is preserved and a replacement is taken from the same popularity-band and category stratum. Version 1.0.0 contains exactly 50,000 unique domains.

This design improves category coverage without pretending that every category occurs equally often. The complete allocation for every band and primary category is supplied in `dataset-analysis.json`.

Measured results

The top 1,000 are more frequently multi-category

Popular platforms often combine communication, publishing, streaming, commerce, hosting and other functions. The measured difference between the two sample components makes that visible.

51%

Top 1,000: 509 multi-category domains

50.9% have more than one category, with an average of 2.191 labels per domain.

41%

Stratified 49,000: 20,292 multi-category domains

41.41% have more than one category, with an average of 1.575 labels per domain.

Most frequent labels in the top 1,000

These are label occurrences, not mutually exclusive groups. One domain can contribute to several bars. The counts describe the popularity-oriented component only and should not be used as internet-wide prevalence estimates.

News
189
Arts
159
Adult
144
Political
104
Shopping
103
Business
97
Sports
96
Streaming
84

What can be learned from the difference?

The result supports a practical point about domain categorization: well-known domains are not necessarily simpler to classify. A single platform may expose several products, host user-generated content, operate an advertising system, distribute media and provide developer infrastructure. Multi-label data preserves those overlapping functions.

The stratified component still contains substantial multi-label behavior, but it also reaches specialized sites whose purpose is narrower. Researchers can compare the two components directly because `sample_segment` and `popularity_band` remain attached to every record.

The complete sample observes 58 of the taxonomy's 59 categories. The included taxonomy files define all 59. A category missing from a finite sample has not disappeared from the full database; it simply was not present among the selected rows.

Classification model

One domain can carry several useful labels

The release preserves the full ordered labels and also provides one primary label for sampling and grouped analysis.

Domain

A normalized domain is the unit of observation.

Ordered categories

All assigned filtering categories are retained.

Research use

Analyze, group, join, benchmark or map categories into a policy vocabulary.

The Web Filtering Database taxonomy contains 59 categories designed for content filtering and policy enforcement. Subject-oriented labels cover areas such as Business, Education/Reference, Health & Medicine, News/Media, Shopping, Sports & Recreation and Travel. Policy-sensitive categories include Adult, Gambling, Illegal Drugs, Violence and Weapons. Operational categories include AI / LLM Tools, CDN / Hosting / Infrastructure, File Sharing/P2P, Malware / Phishing / Suspicious, Remote Access Tools, URL Shorteners and VPN / Proxy / Anonymizers.

The complete category reference, definitions and example domains are available on the Web Filtering Database taxonomy page. That page also publishes a crosswalk to the IAB Content Taxonomy. The systems have different purposes: Web Filtering categories are designed for policy enforcement, while IAB provides an advertising and content-description vocabulary. This sample contains the Web Filtering labels; the crosswalk is included as supporting research material.

Data dictionary

Seven fields, with sampling context kept in every row

CSV and newline-delimited JSON distributions contain the same records and fields.

FieldMeaningInterpretation note
domainLowercase domain from the source snapshot.Unique within the released sample.
categoriesAll ordered Web Filtering category labels.Multiple CSV labels are separated by |.
primary_categoryThe first label in the ordered category list.Used to define mutually exclusive sampling strata.
sample_segmenttop_1000 or stratified_49000.Keep this field when comparing or modeling the data.
popularity_bandThe source-rank interval associated with the record.Supports band-specific analysis.
source_rankPosition in the popularity-ordered source extract.Not an external traffic count or third-party rank.
classification_versionVersion recorded in the source row.All v1.0.0 rows use version_2026_03_15.
Practical uses

Built for analysis, teaching and integration work

The release provides enough real structure to support useful work without presenting a static sample as a production filtering service.

Pipeline testing

Test CSV and JSONL ingestion, multi-label parsing, domain joins, lookup structures and export procedures before connecting a live product feed.

Distribution analysis

Compare category and TLD patterns between recognizable domains and the six stratified popularity bands.

Teaching

Demonstrate stratified sampling, reservoir sampling, multi-label data and the limits of drawing prevalence conclusions from designed samples.

Interface prototypes

Build category explorers, policy editors, dashboards and search interfaces using realistic data and a documented taxonomy.

Taxonomy mapping

Map the filtering vocabulary into an internal policy system or study the supplied crosswalk to the IAB Content Taxonomy.

DNS-log enrichment

Exercise joins between a controlled DNS-log sample and static domain-category data without depending on a live endpoint.

Working with the release

Keep the sampling design attached to every result

The most useful analyses begin by deciding whether the question concerns familiar high-visibility domains, the broader stratified collection, or a comparison between them.

Analyze the components separately first

Start by grouping on sample_segment. Report category frequencies, multi-label rates and TLD distributions independently for the top 1,000 and stratified 49,000 before calculating a combined result. That simple step prevents the deliberately included popular-domain component from being mistaken for a random part of the long tail.

For comparisons across visibility levels, retain popularity_band and use the exact band boundaries documented in the methodology. Researchers can then ask whether multi-label assignments, particular categories or particular TLDs become more or less frequent across the source ordering.

Distinguish labels from organizational policy

A taxonomy states what type of content or function a domain is associated with. It does not prescribe one universal action. A school may restrict a category that a university research lab permits; a financial institution may monitor a service that a home network allows without additional controls.

When prototyping a policy interface, preserve that distinction in the data model. Store the category assignment separately from the organization's allow, block or monitor decision. This makes it possible to revise policy without rewriting the underlying classification and to explain why the same domain receives different treatment in different settings.

Recommended citation practiceState the dataset version, the component or bands analyzed, whether all labels or only primary_category were used, and the date of access. After the Zenodo release is live, include its DOI so another researcher can retrieve the same version.
Limitations

What this dataset should—and should not—be used for

Research value depends on keeping the boundaries of the release clear.

It is a snapshot

Domains change ownership, content and function. Category assignments can become outdated after publication. The source classification version is recorded so later work can identify the exact snapshot used.

It is domain-level

A domain can contain many sections and services. This release does not claim that every individual URL under a multi-purpose domain has identical content or policy relevance.

It is a designed sample

The rank bands and protected category allocation improve evaluation coverage. Unweighted totals describe this research collection, not the complete source database or all active domains.

It is not a production feed

Do not deploy it as the sole blocklist for a school, enterprise, ISP, DNS resolver, firewall, parental-control product or security service. Production use requires broader coverage and ongoing updates.

Why not publish a simple random sample?

A simple random sample would be valid for some statistical questions, but it would contain many unfamiliar long-tail domains and could provide few examples from small categories. This release is intended for evaluation as well as analysis, so it separates recognizable top domains from a broader stratified component.

Why are there 58 observed categories when the taxonomy has 59?

The complete taxonomy has 59 categories. A finite sample does not necessarily contain every category. The included taxonomy files define all 59, while the released domain rows happen to observe 58.

Can categories be converted directly into allow and block decisions?

Categories are inputs to policy. An organization still decides which categories to allow, block or monitor for its users and requirements. The same category can be handled differently in a school, a bank, a home network or a research environment.

How does this differ from the complete database?

The public release contains 50,000 static records. The commercial database currently provides packages covering up to 120 million domains, scheduled refreshes, API and offline delivery, and licensing for internal, on-premise and OEM deployments.

Download the documented research release

The Zenodo package contains CSV and JSONL data, taxonomy and IAB mapping files, the complete methodology, descriptive analysis, source code and SHA-256 checksums.

DOI and download link will appear here when the Zenodo record is published.