The audio is encrypted. The call record is not. This is the metadata surveillance problem applied to real-time communication. Your carrier knows you called a journalist at 11 PM. The platform's signaling server knows the duration. The NAT-traversal relay knows both IP addresses. Every piece of metadata the intelligence analyst needs to build a case exists in infrastructure logs — even if no one ever decrypts a single audio packet.
Most encrypted-voice implementations follow the same pattern: set up a call through a signaling server, negotiate keys, encrypt the media stream with SRTP, tear down the call. The media stream is protected. Everything around the media stream is not. The signaling server knows who called whom. The NAT-traversal servers know the endpoints. The metadata — call duration, time of day, peer identity — is available to the infrastructure even when the audio is perfectly encrypted.
Zentalk's real-time layer treats the call, not just the media, as the unit of protection. SRTP remains the transport cipher, but it sits inside a larger construction that hides signaling, routing, peer identity, and call metadata with the same rigor applied to text messaging. This article walks through the layers and what each one denies to an adversary.
What SRTP actually does
To understand what is missing from the standard approach, start with what SRTP genuinely provides. Secure Real-Time Transport Protocol, standardized in RFC 3711, encrypts and authenticates RTP media packets. A call using SRTP produces ciphertext where the audio or video payload would have been, with integrity tags that prevent tampering. Classical SRTP uses AES-CTR for the cipher and HMAC-SHA1 for authentication, though modern deployments use AES-GCM for combined authenticated encryption.
SRTP is a mature, well-analyzed standard. Its properties are well-understood. What it does not do: it does not negotiate the keys (that is SDP, DTLS-SRTP, or ZRTP's job), it does not route the packets (that is ICE, STUN, and TURN's job), and it does not conceal the peer identities (that is not a problem SRTP was designed to address).
The signaling problem
A voice call over most platforms starts by contacting a signaling server: "I want to call Alice." The server looks up Alice, notifies her, and coordinates the rendezvous. The server sees both endpoints, the call timing, and often the duration. Even if the media is end-to-end encrypted afterward, the call graph is fully visible to the signaling operator.
Zentalk removes the centralized signaling server. Call setup uses the same validator network that routes text messages. A call invitation is an end-to-end encrypted message sent to the callee's routing hash. The callee's client receives it, accepts or declines, and responds with routing information for the media path. The validator network forwards the invitation blindly, without learning who is calling whom.
The call graph — who is on a call with whom, when, and for how long — is not accessible to any single infrastructure component, because no component holds it.
Media path privacy
Zentalk supports real-time calls with up to 6 participants alongside screen sharing — a deliberate ceiling that keeps the encrypted media construction tractable without degrading to a broadcast model. Each participant's media stream carries its own SRTP session key negotiated through the post-quantum hybrid exchange described below, so the 6-participant limit is not a shortcut but a commitment: every seat in the call receives the full privacy treatment.
SRTP protects what flows through the media path. The path itself still needs to be established. For most apps, the path is a direct peer-to-peer UDP flow between the two endpoints, with TURN relays fallback when NAT traversal fails. Either way, the endpoints of the media path are IP addresses that reveal both parties' network locations.
Zentalk offers two modes. Direct mode uses standard WebRTC-style peer-to-peer media with SRTP, accepting that the peers see each other's IP addresses. This mode matches the quality and latency of mainstream video calls and is appropriate when IP-level privacy is not critical.
Routed mode sends the SRTP stream through validator relays using the routing construction used for text messaging. Each packet passes through three independent hops, with RSA-4096 layers that prevent any single relay from learning both endpoints. Latency increases modestly — typically 30 to 80 milliseconds of additional round-trip time — in exchange for the IP-level unlinkability.
Key negotiation
The media encryption key is negotiated through the same post-quantum key exchange construction used for text. X25519 and Kyber-768 key exchanges run in parallel, and the shared secret is derived from both. The derived key feeds SRTP as its session key. The result is a media stream whose confidentiality survives both classical and quantum cryptanalysis, consistent with the rest of the Zentalk protocol.
Keys rotate during the call at short intervals. A single long call does not use a single session key end-to-end; forward secrecy applies to the media layer, so a mid-call key compromise does not expose earlier media segments.
Identity and verification
A call in Zentalk is placed to a wallet-based identity, not a phone number. There is no cellular routing involved, no caller ID service, no central directory mapping numbers to names. The callee sees a cryptographic identity they have previously verified with the caller. An adversary cannot spoof the caller by hijacking a phone number, because phone numbers are not part of the identity system.
Out-of-band verification — a fingerprint exchange or QR scan — establishes the identity binding once, and the binding persists across calls. The user sees a verified indicator for contacts whose fingerprints they have checked, and a warning for contacts whose identity is new or unverified.
Metadata beyond the call
Call history, missed-call notifications, and voicemail-like features are handled as encrypted messages, subject to the same per-message forward secrecy and metadata minimization as regular text. A record that a call occurred exists on the participating devices; it does not exist on any server, and it does not exist in a form that can be extracted by any third party.
Contact-list access for calling is explicit. The user selects a contact from their own address book (stored locally on the device) and initiates the call. No upload of contacts to a server is required or performed.
What this does not solve
A voice call exposes the caller's voice to the callee. If the callee records, transcribes, or analyzes the audio, Zentalk's cryptography does not prevent it. The media is encrypted between the endpoints, not past them. The user is trusting the person they called with the content of what they said.
Similarly, ambient audio captured by the microphone (background conversations, environmental sounds) is part of the stream. Zentalk does not filter what the microphone picks up. Users with specific privacy concerns about the environment around a call should manage that at the physical level.
Why this matters
For most users, a voice call is the most intimate mode of communication they use. More intimate than messaging. More spontaneous than email. And, on most platforms, the mode where the metadata disclosure is most complete: a record of who called whom, when, and for how long, retained by carriers for years.
Zentalk's call layer brings the same threat model to real-time communication that the text layer brings to asynchronous messaging. The media is encrypted. The signaling is decentralized. The call graph is not stored. The identity is self-sovereign. A call is a conversation that is private in the way the word is supposed to mean, rather than a conversation that is mostly private except for the metadata trail.
Thanks & Best Regards Zentachain Team!



