Persistence Blog

Outbound at Scale: How Persistence Powers 100,000+ Simultaneous Calls

7 min readJUL 14, 2026

What "scale" actually means for outbound voice AI

Most voice AI platforms demo with 10, 100, maybe 1,000 concurrent calls. Real enterprise outbound campaigns — insurance renewal notices, debt collection, appointment confirmations, political surveys, retail promotions — run at 10,000–100,000 simultaneous calls. At that volume, any platform weakness becomes a critical failure. Latency spikes that are invisible at 100 calls cause cascading degradation at 50,000. API rate limits that never trigger in testing block calls in production. Telephony infrastructure that works in the US silently fails for international dialing. Persistence is the only platform purpose-built for this scale from the ground up.

Telephony infrastructure: what's under the hood

Persistence operates direct carrier interconnects in the US, EU, and APAC — we are not reselling another provider's SIP trunking. This direct carrier relationship gives us three advantages: lower per-minute costs that we pass to customers, higher call completion rates because our numbers carry better carrier reputation scores, and priority handling during carrier congestion events. Bland AI and Retell AI both use commodity SIP trunking providers, which means their call quality and completion rates are dependent on that provider's infrastructure decisions and capacity.

Spam labeling: the silent call killer

Outbound AI calling at scale runs into a serious problem that most platforms don't discuss: spam labeling. Mobile carriers and third-party services like STIR/SHAKEN and First Orion flag numbers with unusual call patterns as "Spam Likely," causing a significant percentage of outbound calls to be rejected or ignored. Persistence manages a large pool of verified, warmed phone numbers with call pattern diversity algorithms that keep numbers out of spam databases. We actively monitor number reputation across all major carriers and rotate flagged numbers automatically. The result: our customers see 15–25% higher answer rates on outbound campaigns compared to competitors running off commodity number pools.

Batch calling, scheduling, and rate limiting

Persistence's outbound orchestration layer gives you fine-grained control over call scheduling. You can define call windows by timezone (never call before 8am or after 9pm local time), set per-campaign concurrency limits, and configure retry logic with customizable backoff curves. Our batch calling API accepts lists of up to 500,000 contacts and distributes calls across your configured time windows automatically. You can monitor campaign progress in real time, pause and resume calls, and pull granular analytics — connection rate, transfer rate, call outcome distribution — at any point during the campaign.

Post-call analytics at scale

At 100,000 calls, you cannot listen to individual recordings to understand what's working. Persistence's post-call analysis pipeline runs on every call automatically: it extracts call outcomes (appointment scheduled, objection raised, callback requested), classifies sentiment, flags calls where the AI failed to handle a question, and surfaces these insights in an aggregate dashboard. You can drill into any call from the dashboard, read a full transcript, listen to audio with timestamp navigation, and mark calls for agent retraining. This closed loop between production calls and model improvement is what keeps Persistence agents getting better over time.