Key Takeaways
- Bulk verification means processing a large list efficiently: stream the CSV, cap concurrency, handle retries, and write results incrementally.
- Bounded concurrency with a semaphore keeps you within the API rate limit while still verifying many addresses in parallel.
- Transient failures and rate-limit responses should trigger a retry with backoff, not a discarded row.
- Write results as you go and sort by verdict, so a crash never loses progress and the output is ready to action.
Verifying one address is a single API call. Verifying fifty thousand is an engineering problem: you need to stay within rate limits, survive transient failures, avoid losing progress if the job dies, and end up with output you can act on. Bulk email verification is the workflow that turns a raw list into a clean, sorted one at scale. This guide builds a production CSV-to-CSV pipeline in Python against the v2 API, with bounded concurrency and retry handling. The same patterns apply in any language; see the email verification integrations hub for other runtimes.
The four principles that make bulk verification reliable are: stream rather than load everything into memory, bound concurrency, retry transients, and checkpoint results incrementally.
The v2 Endpoint
Each address is verified with a single GET request. Confirm the shape of the response before writing the batch loop.
For bulk work the fields that matter most are status (passed, failed, unknown, transient) and the risk flags, which together decide which output file each row belongs in.
The Bulk Verifier
This Python worker uses asyncio and aiohttp to verify addresses concurrently, bounded by a semaphore so no more than a fixed number of requests are in flight at once. It reads from a CSV and writes results as they complete.
Three details make this production-grade. The semaphore caps concurrency at 20 regardless of list size, keeping you within the rate limit. A 429 or a transient status triggers exponential backoff and a retry rather than discarding the row. And flushing after every write means a crash at row 40,000 leaves you 40,000 verified rows, not an empty file.
Sorting Results by Verdict
A flat results file is not actionable; you want it split by decision. This step reads the results and routes each row to valid, invalid, or risky, matching the send-or-suppress logic you will apply downstream.
Now you have valid.csv to send, invalid.csv to suppress permanently, and risky.csv to review by flag. That is the whole point of bulk verification: a raw list in, three actionable lists out.
One refinement worth adding for very large jobs is resumability. Because the verifier checkpoints every row, a second run can read the existing results file first, build a set of already-verified addresses, and skip them before starting. That turns a crashed 200,000-address job into a quick continuation rather than a full restart, and it also lets you re-run periodically while only paying to verify the addresses that are genuinely new since last time.
New accounts can run a real batch against their own list with 100 free email verification credits, and the email verification API documentation covers rate limits and the full response schema. Python developers can find the language-specific reference on the verify email with Python page.
Frequently Asked Questions
How do I verify a large email list without hitting rate limits?
Bound your concurrency with a semaphore so a fixed number of requests are in flight at once, and handle 429 responses with exponential backoff and a retry. Start with a moderate concurrency value, watch for rate-limit responses, and increase only if you are staying under the limit. This keeps a large batch moving without triggering throttling.
What happens if my bulk verification job crashes partway through?
If you write results incrementally and flush after every row, a crash leaves you with all the rows verified so far rather than an empty file. You can then resume from where you left off by skipping already-verified addresses. Checkpointing is the single most important habit for large jobs.
Should I deduplicate my list before verifying it?
Yes. Lists assembled from multiple sources frequently contain the same address several times, and verifying duplicates wastes credits and time. A simple deduplication step before the batch loop, using a set, can meaningfully cut the number of API calls and the cost of the run.
How should I organize the output of a bulk verification run?
Sort results by verdict into separate outputs: valid addresses to send to, invalid addresses to suppress permanently, and risky addresses to review by flag. Splitting the flat results file by decision makes the output immediately actionable and mirrors the send-or-suppress logic you apply in your email platform.