toward a multimedia telephone network
The FCC’s Wireline Competition Bureau closed the first day of its IP Transition Workshop on July 15, 2026 with a panel on consumer protection, co-moderated by Chris Laughlin of the Competition Policy Division and Michael Scott of the Consumer and Governmental Affairs Bureau’s Disability Rights Office. I was on it, with David Bahar of Telecommunications Access of Maryland, Rebekah Johnson of Numeracle, Joshua Ruby of Granite, and Jim Tyrell of TNS. The co-moderation was deliberate — the Bureau put consumer protection and accessibility in one conversation instead of two — and the part of my contribution I care about most came out of that pairing. I argued that an all-IP network can make real-time text, video relay, multimedia communication and NG911 native capabilities of the telephone network rather than bolt-ons grafted onto a voice path that was never designed for them. I wrote up the two days in a separate entry. This is the argument at length.
For most of its history the telephone network did one thing: it carried voice, in real time, between two numbers. Everything else we ever wanted from it — text for people who cannot hear, video for people who sign, images and data on an emergency call, a verified name for the person calling — got added afterward, in a separate system, on a separate island, reachable only if you knew the trick. The all-IP transition is the first real chance to stop bolting things on. If we build it deliberately, the telephone network can become natively multimedia: voice, real-time text, video, and rich data as first-class properties of a call, addressed by a telephone number, carried end to end, and reachable by anyone from anywhere.
I want to make the case for that as a goal — not a product, but the thing the transition is for.
The pieces already exist
The striking thing, once you look, is that almost every component of a multimedia telephone network already exists as a standard. Real-time text is defined to ride inside the same IP media session as voice — a text stream negotiated alongside audio, character by character, in the “total conversation” model. Video relay is number-addressed and interoperable in principle, with a shared numbering directory and profiles for both providers and user equipment. NG911 can already carry text, images, and video to a call center over IP. And the identity work I have spent years on — STIR/SHAKEN, rich call data, verified caller name — lets identity and context travel with a signed call rather than being looked up separately at the far end.
The problem is not that these do not exist. It is that they exist in parallel, half-connected, each on its own island. Real-time text works, but fractures at carrier boundaries and disappears the moment you leave the wireless network. Video relay works, but lives in a numbering and identity plane separate from the rest of the network. NG911 multimedia works where the emergency system has been upgraded, which is unevenly. Verified identity is signed between carriers but does not reach the PSAP. Each capability is real; none of them is native.
What “native” means
A capability is native when it is a property of the network rather than an application on top of it. Native means you do not need a special app, a special number, a special provider, or foreknowledge that the person you are calling needs a particular mode. You place a call to a telephone number, and the call negotiates what it can carry — voice, text, video, the caller’s verified identity, the rich data that gives a call context — and both ends use whatever they need. The telephone number stays what it has always been: the way people identify themselves and reach each other. What changes is that the number now addresses a multimedia session, not just a voice circuit.
This is the difference between a network that permits text and video through gateways and overlays, and one where text and video are simply media the call can carry, the way voice always has been.
Design to the calls that matter most, not the lowest common denominator
There is a habit in network design of building to the lowest common denominator — the minimum every endpoint is guaranteed to support — and treating everything above it as optional. For the telephone network I think that instinct is backwards. The capabilities that should define the network are the ones that matter most when they matter at all: reaching 911, and reaching someone who depends on an accessible mode to communicate.
Not everyone needs 911 on any given day. But when someone needs it, it has to work — and it has to carry everything that helps: location, identity, and, increasingly, text and video, because a caller who cannot speak safely, or at all, still has to be understood. Those capabilities are already sitting in nearly every mobile device we carry. If the handset can do real-time text and video out of the box, and the emergency system is being rebuilt on IP anyway, then building the emergency call to use them is not a stretch goal — it is the obvious thing.
The same logic runs through accessibility. Not everyone needs real-time text or video relay. But for the people who depend on them they are the difference between having a phone and not, and for everyone else they are simply a richer way to communicate — again, using functionality already built into the devices in our pockets. Faced with that, the question people reach for is “why not.” But “why not” almost cheapens it. The honest question is why didn’t we — why, with the pieces already in hand, we left these as bolt-ons and gateways instead of building them in from the start. The reasons are real and many: separate systems, separate funding, separate standards efforts, the inertia of a network that grew one capability at a time. None of them is a reason we cannot correct it now, especially once we get to a unified all-IP network and agree on the standards going forward.
There is a useful test buried in this. Accessibility is the honest diagnostic of whether the network is really multimedia: if real-time text, video relay, and NG911 are genuinely native — reachable by number, carried end to end, working from any device on any network — then the network is natively multimedia, because those are exactly the modes that exercise every part of it. If they are still bolted on, reachable only through a particular provider or a gateway that downgrades the experience, then it is still a voice network with attachments, whatever the marketing says. Design to the most critical and most dependent users, and you raise the floor for everyone — the opposite of lowest-common-denominator thinking, and the right way to decide what an all-IP telephone network should be able to do.
Identity is what makes it trustworthy
Multimedia without identity is not an improvement — it is a larger attack surface. The reason the identity work and the multimedia vision belong in the same sentence is that a network carrying more — text, video, branded call data, a caller’s name — has to carry proof along with it, or the added richness just gives bad actors more to spoof. Trust becomes transitive when it travels with the identifier: a call that is cryptographically bound to a telephone number, and to the entity authorized to use it, can carry verified context that the far end can actually rely on. The multimedia telephone network and the trustworthy telephone network are the same project. You cannot have the first without the second.
The honest gaps
I do not want to pretend this is close. The distance between “the standard is done” and “the capability is native” is exactly where real users get stranded, and it is wide. Real-time text still breaks across carriers, handset makers, and operating systems, and is nearly absent on wireline and in over-the-top apps. Video relay is still siloed in proprietary platforms and a separate numbering plane. Verified identity is not deployed to PSAPs, so the most important call a person ever makes is also the one where the network knows the least about who is on the line. Each of these is a solvable engineering-and-policy problem, and none of them is solved.
But the transition is the moment these become solvable at all, because it is the moment the underlying network stops being the constraint. The reasons these capabilities are half-built are mostly the seams of the old network — the TDM hops, the legacy gateways, the provider islands. Take those away, and what is left is the work of making the standards we already have interoperate end to end.
The record has since said it in filings
In the five weeks after the workshop, the diagnostic above showed up in the Commission’s own record.
On August 10, four telecommunications relay providers answered paragraphs 123 and 124 of the know-your-upstream-provider notice, which asked how STIR/SHAKEN applies to TRS. InnoCaption told the Commission that TRS providers “cannot obtain Service Provider Code (‘SPC’) tokens, as they do not meet the STIR/SHAKEN Governance Authority’s requirements to obtain a token.” Hamilton Relay wrote that for PSTN-based relay “the conferenced nature of these calls means they cannot satisfy the criteria for A-level attestation.” ZP Better Together quoted the Commission’s own notice back at it: downstream filtering of unsigned video relay calls “disproportionately harm[s] individuals with disabilities.” Sorenson and CaptionCall supplied a fix, observing that A-level attestation “consists of three components that can be performed by different providers” and asking that originating providers be required to recognize a certificate carrying the relay provider’s own attestation of end-user verification. Four providers, none of them able to hold the credential, on calls Congress mandated under section 225. (I co-chair the ATIS/SIP Forum IP-NNI Task Force and the STI-GA Technical Committee, which is where the token requirements those filings describe are written.)
On August 13, the Wireline Competition Bureau filed the workshop letter and the full transcript into six dockets — WC 26-96, 25-311, 25-304, 25-208, 10-90, and 17-97, the call authentication proceeding. The panel discussion is now part of the authentication record rather than an event that happened near it.
These filings are the diagnostic stated as a complaint. Relay is the case where the bolt-on is not a matter of user experience: a service Congress mandated under section 225 arrives unsigned, and the network’s own defenses against fraudulent traffic then treat it accordingly. The accessible mode is not merely gatewayed, it is gatewayed in a way that makes it look less trustworthy than the traffic it is competing with for delivery.
The fix Sorenson names is the one this argument has been pointing at throughout. If a capability is native, the identity that goes with it is native too — a certificate the relay provider holds and presents, verified by whoever receives the call, rather than an attestation the relay provider is structurally unable to make. That is the same move as making real-time text a media stream rather than an app: put the thing in the call instead of beside it.
What is missing is a way to tell whether any of it happened. The transition has cost milestones and date milestones and no milestone about what the network still has to do when it is finished. Name the properties that have to survive — a relay call that can be signed, a signature that arrives at termination, an accessible mode reachable by number from any device — and require that somebody report whether they do. Without that, every filer can cite accessibility or call authentication as a benefit of finishing the transition and nobody owes a number.
What we could strive for
So here is the goal I would put on the wall: a single telephone network, all-IP, where a call to a number can carry voice, real-time text, and video as ordinary media; where the caller’s identity and the call’s context travel with it, verified; where the accessible modes are native rather than gatewayed; and where all of it is reachable from any device, anywhere, without a special app or a special provider. Not a new network — the telephone network, finally built to do what we have always wanted it to do.
The transition off TDM will happen regardless; the only question is whether we treat it as a cost-cutting exercise or as the once-in-a-generation chance to build the network we would design if we were starting today. I know which one I want to spend my time on.