Krish Kumar, CEO at Wowza.
getty
Picture an autonomous surveillance drone, 200 kilometers out over open ocean. Past 150 kilometers, its satellite link drops to roughly 50 kilobits per second—not enough bandwidth to stream a single second of usable video. There is no cloud out there; there is no control room watching. If that drone is going to detect a vessel where it shouldn’t be, it has to see, decide and act entirely on its own.
I opened with one of the hardest versions of the problem on purpose, because a milder version of it is sitting inside almost every organization I talk to. The last few years of AI progress have been built on the assumption that data can travel to where the intelligence is. For text, that assumption holds—a prompt is a few kilobytes. For video, the physics, the security rules and the economics get in the way.
The Reasons Video Can’t Travel
Three things keep operational video out of the cloud: security, connectivity and cost.
1. Security
In many enterprise environments, video is simply not permitted to leave the building. Hospitals run cameras in intensive care units. Courts and corrections facilities record interview rooms. Defense and government installations operate on classified networks that never touch the public internet. Critical-infrastructure operators run air-gapped systems by design. For these organizations, footage leaving the network is not a policy inconvenience—it is a breach.
2. Connectivity
Some video physically cannot travel because there is no connection large enough to carry it. This includes an offshore rig on a constrained satellite link, a research facility with cameras running around the clock in underground caverns, or a fleet of vehicles or drones operating far beyond any network built for continuous high-bitrate uplink.
3. Cost
One camera at a modest 2 megabits per second generates more than 20 gigabytes a day. Multiply that across hundreds or thousands of cameras, and the bill for centralized analysis scales faster than almost any other data problem in the enterprise.
Why The Cloud Became The Default
For two decades, the safe answer to most infrastructure questions has been the cloud. Elastic scale, no hardware to maintain and paying for what you use made it the logical destination for enterprise workloads. Video followed the same path as everything else, and for storage and delivery, it mostly made sense.
AI deepened the dependency. The models that make video analysis credible run on data center GPUs, and most commercial video AI offerings are built around a cloud model: upload footage, receive detections or labels. This works well as long as your video can reach them (and is allowed to reach them) and the bill stays manageable.
Moving The Intelligence To The Edge
Increasingly, I’ve been hearing a different request from customers: “Help us run AI without the cloud.” The reason this is now possible and practical comes down to the fact that detection is not footage.
“Vessel entering restricted zone, camera 7, 14:03:22” is a few hundred bytes. A clip of the moment that matters is a few megabytes. The raw stream is terabytes. When the model runs where the camera is—on the vessel, on the rig and inside the secure facility—the only thing that needs to travel is the answer. That drone on a 50-kilobit link cannot send video to anyone. It can absolutely transmit a line of metadata and act on what it finds.
I see this across our own customer base. A hospital analyzes patient-room video for falls and distress without a single frame leaving its network. A drilling operator watches for red-zone incursions on the rig itself and sends shore only the events, not the footage. A transportation agency runs detection on cameras in the field rather than hauling thousands of streams back to a data center. A drone at the edge of its link completes its mission instead of abandoning it.
A detection is not a report to be filed. It is a trigger for whatever should happen next. A drone might change course to investigate, a rig may sound an alarm on its local network and a hospital can page the nurses’ station down the hall. None of this requires new cameras. It requires deciding that the intelligence belongs with the video, wherever the video happens to live.
The Question That Should Come First
In fully air-gapped environments, the entire loop—camera, model and response—runs inside the boundary, and nothing leaves at all. When operators do need centralized visibility, they can send events rather than footage. A few hundred bytes of metadata can traverse links that could never carry continuous video, while still respecting the security and governance requirements that keep the underlying footage local.
Regulated industries are beginning to ask not just what a model costs, but where it runs and who controls it. “AI sovereignty” has moved from a policy talking point to a line item in enterprise evaluations, and I believe video pushes that logic to its endpoint, because video is the heaviest data to move and the most sensitive when it moves.
For anyone evaluating video AI, one question belongs at the top of the list, before accuracy and before features: Where is this allowed to run? If the same detection can’t run in the cloud for one facility, on-premises for another and on a device in the field for a third, you risk choosing a system that struggles to meet the realities of production.
The first wave of AI taught organizations to bring their data to the intelligence. The next wave runs in the opposite direction. The organizations that recognize this shift early will be the ones that turn cameras from passive recording devices into active participants in operational decision-making.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?











