Match your domain list against a large gzip file
Find exact names shared by your list and a domain or subdomain dataset. Try the complete workflow with free samples first; the large file is streamed rather than loaded into memory.
- Free inputs
- 2 samples
- Match type
- Exact
- Large file
- Streamed
1,000 domains and 1,000 subdomains
after lowercasing and whitespace trimming
your own list is held in RAM
This tutorial uses the free GetDomainLists samples. They contain historically observed names, not a current DNS or registration check. You do not need to buy anything to run the example.
Choose the right file
Use the free domains sample for registrable names such as example.com or example.co.uk. Use the free subdomains sample for names below those roots, such as www.example.com or api.example.co.uk. These are separate, non-overlapping file types. A host will not match merely because its root appears in the domains file.
Prepare your input before matching
Put one hostname or registrable domain on each line. The public Python toolkit lowercases and trims whitespace, but it does not parse URLs, remove a trailing dot, or convert Unicode IDNs. Extract the hostname from any URL, remove its path, port and trailing dot, and convert internationalized names to ASCII xn-- A-labels with an IDNA-aware tool before matching. For example, https://www.example.com/path is a URL, not a matchable name; www.example.com belongs in the subdomains lookup. Keep your original record ID separately if you need to join results back to your own table.
Run the free example
Python 3 standard library is enough. The public repository also keeps these as data/domains-sample.txt and data/subdomains-sample.txt. Download a dated sample, take two names known to be present, and add one deliberate miss:
git clone https://github.com/getdomainlists/dataset.git
cd dataset
curl -fsSL https://getdomainlists.com/samples/getdomainlists-domains-sample.txt -o domains-sample.txt
head -n 2 domains-sample.txt > my-names.txt
printf '%s\n' 'not-in-sample.invalid' >> my-names.txt
python3 scripts/match_names.py my-names.txt domains-sample.txt > hits.txt
wc -l hits.txt
hits.txt should have 2 lines. The script also prints a 2 of 3 summary to stderr. Repeat with the subdomains sample when your inputs are hostnames. The curated samples are shuffled and are not statistically representative of the full files.
Use a larger file without loading it all into RAM
The same script accepts .txt.gz as the dataset argument: python3 scripts/match_names.py my-names.txt your-domains.txt.gz > hits.txt. It loads your list into a set and streams the dataset line by line; RAM therefore grows with your list and its matches. A gzip scan is sequential, so expect work proportional to the size of the file. If your own input is also too large for RAM, this script is not the right tool; use a separately tested external-sort/merge workflow. Do not run comm directly on these shuffled samples.
For adjacent jobs, read_sample.py previews and counts a plain-text or gzip list, while filter_tld.py streams names for one TLD. The GitHub README has their commands and data terms.
Understand the result
A hit means the normalized exact name appears in the chosen historical file. A miss may be caused by a different file type, a URL/trailing-dot/IDN formatting mismatch, or incomplete dataset coverage. Neither outcome proves whether a name resolves, is registered now, hosts a website, or belongs to a particular company. This is not fuzzy matching or live discovery.
Try the format first with the free 1,000-name samples. Both full files are $9 once, with 30 days to download.
For data patterns rather than matching instructions, see how subdomains are distributed or the dated counts for every TLD.