Inspectable
The first segment contains domains people are likely to recognize. Researchers can examine concrete classifications instead of starting with thousands of unfamiliar long-tail names.
A documented, reproducible sample from the Web Filtering Database: 1,000 popularity-ranked domains for recognizable inspection and 49,000 domains sampled across category and popularity strata.
The release is large enough for analysis and software testing, but small and static enough to remain clearly separate from the complete production database.
The first segment contains domains people are likely to recognize. Researchers can examine concrete classifications instead of starting with thousands of unfamiliar long-tail names.
The second segment reaches across six source-rank bands and primary categories, providing substantially more variety than a list made only from the most popular sites.
The package includes the sampling program, fixed random seed, allocation table, machine-readable analysis, checksums, taxonomy files and a complete data dictionary.
If a sample contains only famous domains, it is easy to understand but biased toward large platforms and heavily visited content. Those domains are also more likely to support several functions from the same destination. If a sample is drawn uniformly from millions of records, it reaches the long tail but can become difficult to inspect and may provide too few examples of small categories.
This release does not hide that tradeoff. The first 1,000 unique domains follow the popularity order of the source extract. The remaining 49,000 are selected separately across popularity bands and primary-category strata. Every row states which component and band it came from.
The commercial database currently covers up to 120 million domains. This release was generated from a 20,128,938-record source snapshot.
Each band is subdivided by the first ordered category label. Smaller available categories receive protected representation before the remaining capacity is allocated proportionally.
Within each rank band, domains are grouped by `primary_category`. The procedure first protects up to 25 observations for each available category stratum. It then distributes the remaining places in proportion to the sizes of the remaining strata and selects rows using reservoir sampling.
The fixed seed is 20260922. Duplicate domains are removed: the earliest selected occurrence is preserved and a replacement is taken from the same popularity-band and category stratum. Version 1.0.0 contains exactly 50,000 unique domains.
This design improves category coverage without pretending that every category occurs equally often. The complete allocation for every band and primary category is supplied in `dataset-analysis.json`.
Popular platforms often combine communication, publishing, streaming, commerce, hosting and other functions. The measured difference between the two sample components makes that visible.
50.9% have more than one category, with an average of 2.191 labels per domain.
41.41% have more than one category, with an average of 1.575 labels per domain.
These are label occurrences, not mutually exclusive groups. One domain can contribute to several bars. The counts describe the popularity-oriented component only and should not be used as internet-wide prevalence estimates.
The result supports a practical point about domain categorization: well-known domains are not necessarily simpler to classify. A single platform may expose several products, host user-generated content, operate an advertising system, distribute media and provide developer infrastructure. Multi-label data preserves those overlapping functions.
The stratified component still contains substantial multi-label behavior, but it also reaches specialized sites whose purpose is narrower. Researchers can compare the two components directly because `sample_segment` and `popularity_band` remain attached to every record.
The complete sample observes 58 of the taxonomy's 59 categories. The included taxonomy files define all 59. A category missing from a finite sample has not disappeared from the full database; it simply was not present among the selected rows.
The release preserves the full ordered labels and also provides one primary label for sampling and grouped analysis.
A normalized domain is the unit of observation.
All assigned filtering categories are retained.
Analyze, group, join, benchmark or map categories into a policy vocabulary.
The Web Filtering Database taxonomy contains 59 categories designed for content filtering and policy enforcement. Subject-oriented labels cover areas such as Business, Education/Reference, Health & Medicine, News/Media, Shopping, Sports & Recreation and Travel. Policy-sensitive categories include Adult, Gambling, Illegal Drugs, Violence and Weapons. Operational categories include AI / LLM Tools, CDN / Hosting / Infrastructure, File Sharing/P2P, Malware / Phishing / Suspicious, Remote Access Tools, URL Shorteners and VPN / Proxy / Anonymizers.
The complete category reference, definitions and example domains are available on the Web Filtering Database taxonomy page. That page also publishes a crosswalk to the IAB Content Taxonomy. The systems have different purposes: Web Filtering categories are designed for policy enforcement, while IAB provides an advertising and content-description vocabulary. This sample contains the Web Filtering labels; the crosswalk is included as supporting research material.
CSV and newline-delimited JSON distributions contain the same records and fields.
| Field | Meaning | Interpretation note |
|---|---|---|
domain | Lowercase domain from the source snapshot. | Unique within the released sample. |
categories | All ordered Web Filtering category labels. | Multiple CSV labels are separated by |. |
primary_category | The first label in the ordered category list. | Used to define mutually exclusive sampling strata. |
sample_segment | top_1000 or stratified_49000. | Keep this field when comparing or modeling the data. |
popularity_band | The source-rank interval associated with the record. | Supports band-specific analysis. |
source_rank | Position in the popularity-ordered source extract. | Not an external traffic count or third-party rank. |
classification_version | Version recorded in the source row. | All v1.0.0 rows use version_2026_03_15. |
The release provides enough real structure to support useful work without presenting a static sample as a production filtering service.
Test CSV and JSONL ingestion, multi-label parsing, domain joins, lookup structures and export procedures before connecting a live product feed.
Compare category and TLD patterns between recognizable domains and the six stratified popularity bands.
Demonstrate stratified sampling, reservoir sampling, multi-label data and the limits of drawing prevalence conclusions from designed samples.
Build category explorers, policy editors, dashboards and search interfaces using realistic data and a documented taxonomy.
Map the filtering vocabulary into an internal policy system or study the supplied crosswalk to the IAB Content Taxonomy.
Exercise joins between a controlled DNS-log sample and static domain-category data without depending on a live endpoint.
The most useful analyses begin by deciding whether the question concerns familiar high-visibility domains, the broader stratified collection, or a comparison between them.
Start by grouping on sample_segment. Report category frequencies, multi-label rates and TLD distributions independently for the top 1,000 and stratified 49,000 before calculating a combined result. That simple step prevents the deliberately included popular-domain component from being mistaken for a random part of the long tail.
For comparisons across visibility levels, retain popularity_band and use the exact band boundaries documented in the methodology. Researchers can then ask whether multi-label assignments, particular categories or particular TLDs become more or less frequent across the source ordering.
A taxonomy states what type of content or function a domain is associated with. It does not prescribe one universal action. A school may restrict a category that a university research lab permits; a financial institution may monitor a service that a home network allows without additional controls.
When prototyping a policy interface, preserve that distinction in the data model. Store the category assignment separately from the organization's allow, block or monitor decision. This makes it possible to revise policy without rewriting the underlying classification and to explain why the same domain receives different treatment in different settings.
primary_category were used, and the date of access. After the Zenodo release is live, include its DOI so another researcher can retrieve the same version.Research value depends on keeping the boundaries of the release clear.
Domains change ownership, content and function. Category assignments can become outdated after publication. The source classification version is recorded so later work can identify the exact snapshot used.
A domain can contain many sections and services. This release does not claim that every individual URL under a multi-purpose domain has identical content or policy relevance.
The rank bands and protected category allocation improve evaluation coverage. Unweighted totals describe this research collection, not the complete source database or all active domains.
Do not deploy it as the sole blocklist for a school, enterprise, ISP, DNS resolver, firewall, parental-control product or security service. Production use requires broader coverage and ongoing updates.
A simple random sample would be valid for some statistical questions, but it would contain many unfamiliar long-tail domains and could provide few examples from small categories. This release is intended for evaluation as well as analysis, so it separates recognizable top domains from a broader stratified component.
The complete taxonomy has 59 categories. A finite sample does not necessarily contain every category. The included taxonomy files define all 59, while the released domain rows happen to observe 58.
Categories are inputs to policy. An organization still decides which categories to allow, block or monitor for its users and requirements. The same category can be handled differently in a school, a bank, a home network or a research environment.
The public release contains 50,000 static records. The commercial database currently provides packages covering up to 120 million domains, scheduled refreshes, API and offline delivery, and licensing for internal, on-premise and OEM deployments.
The Zenodo package contains CSV and JSONL data, taxonomy and IAB mapping files, the complete methodology, descriptive analysis, source code and SHA-256 checksums.