
About six months ago, our Happy Machines team began researching the camera sensor space and its cost reduction through economies of scale. This cost reduction has had a proliferating effect across every market segment we know. Our initial problem space was the home security market. Greenware Systems, a home security solutions agency, funded the research to build the most cost-effective security camera. We called this Project Chimera, after the Greek mythology creature. It was an apt name, since we could see camera hardware and software forming the two-headed beast.
The whole project seemed straightforward, but we still wanted to step back and start from first principles. Identify correlations and determine if we could build a cost-efficient prototype, both in terms of hardware and software. While doing so, we also sought to understand the evolution of camera sensors and computer vision technologies.
The initial phase of this project involved compiling all our research, during which we began to understand how this also correlates with the evolution of information sharing and digital communication as of today. We're witnessing the emergence of video as the dominant paradigm for information exchange, not because of social media and all the TikTok & reels, but because camera hardware has finally become genuinely accessible. I have compiled the most interesting parts of the research as the following:
Single camera sensor
We started with the most basic element, a single camera sensor. I selected a single sensor, Sony's new IMX611 SPAD (Single-Photon Avalanche Diode) Time-of-Flight depth sensor, which was announced in March 2023, with a sample price of approximately $7-8 USD. That's mind-blowing. Only a few years ago, the cost of storage and camera sensors was several times higher. This doesn't feel like an incremental improvement; it's an exponential shift that's making video capture far more accessible.

The scale of adoption
We began to look further downstream at the effects of adoption and found that an estimated 6.57 billion camera modules were shipped globally. That's not a niche gadget anymore; camera technology has become as ubiquitous as silicon. Amongst most sectors, the automotive industry has been an unexpected driver. Self-driving cars might be the next step in our society's evolution of cars, but automotive ADAS requirements are pushing camera capabilities far beyond what consumer electronics demand.

Take dashcams, one of the most basic uses of camera sensors. They're now available for as little as $50. I remember installing mine for around $500 in 2017 for my new Golf. Although it's not a direct apples-to-apples comparison, we can see the cost of the hardware come down drastically in the last few years.
Custom silicon for processing
Then we shifted focus a bit laterally on how all this data captured by these sensors was being processed. We also observed a move towards custom processing silicon. The volume of data produced by modern sensors is staggering. A 200MP sensor capturing multiple frames in quick succession for an HDR image demands serious computational power. We can see this every year with the release of smartphones with better cameras and computation.
The industry's solution was the widespread adoption of custom silicon, including increasingly powerful Image Signal Processors (ISPs) tightly integrated with AI accelerators, such as Neural Processing Units (NPUs). Google's Tensor chip, Vivo's V1+ imaging co-processor, and Oppo's MariSilicon X are prime examples of this trend.
What I find particularly interesting is how tightly coupled this co-evolution is. The advancement of sensors and processors isn't sequential; they're locked in a symbiotic relationship. Custom silicon is not a marketing gimmick. It's an architectural necessity to get the most out of modern camera hardware.
The 3D sensing breakthrough
Alongside the evolution of 2D imaging, how people consume captured video has also been changing. The appetite to capture and consume through imaging keeps growing. For mass-market consumer devices, our research indicates that the industry is largely converging on active Time of Flight (ToF)/Light Detection and Ranging (LiDAR) technology.
The deliberate choice of active systems over passive stereo vision is becoming clear. As an active system that provides its own illumination, ToF works reliably in the dark and on textureless surfaces, common scenarios where passive stereo vision often fails. This investment has built a large installed base of 3D-aware devices, normalising the concept of 3D capture for millions of consumers.
The software tie up
The evolution of hardware also advances computational photography. Machine learning algorithms have become as important as the physical lens and sensor, governing nearly every aspect of the imaging pipeline.
The most interesting development we tracked was the rapid advancement in Monocular Depth Estimation (MDE). This infers a full, dense 3D depth map from a single 2D image or short video stream, using only one camera. It effectively "dematerialises" the depth sensor, turning it from physical silicon into an algorithm. Soon we will see this affect the way we consume multimedia.
Here is an interesting read:
"Shakes on a Plane," which demonstrated that high-quality, metric-accurate depth could be recovered from tiny, natural parallax created by user hand tremors in a short video burst.
Video as information infrastructure
As camera hardware costs dropped and capabilities grew, we also noticed a shift in how information flows through systems. Video isn't only becoming more prevalent; it will continue to grow as the primary medium for capturing, analysing, and sharing information. From 2023 to 2029, CIS shipments are expected to rise from 6.8 billion to 8.6 billion units. That's billions of sensing endpoints being deployed across all sorts of applications. Each camera becomes a data collection point, turning the physical world into queryable information.

In our research, we documented applications from manufacturing quality control and agricultural monitoring to traffic flow analysis and retail behaviour tracking. The common thread? Video data is becoming the default way to understand dynamic systems. It's not just about recording events. It's about creating a machine-readable layer of reality where every frame carries actionable data.
The democratisation effect
What strikes me most about our research findings is how camera hardware has become genuinely democratised. The barriers to implementing sophisticated visual sensing systems have largely evaporated. When anyone can capture, process, and share high-quality video data, it naturally becomes the go-to format for preserving and transmitting information.
The machine-readable future
The camera is no longer primarily a tool for capturing a static, 2D moment for a human to view later. It's now an ambient, always-on sensor for capturing a dynamic, 4D model of the world, enabling a machine to analyse and understand it.
I think of this as "ambient documentation," where continuous video capture creates searchable and analysable records of everything happening in physical spaces. Video is becoming as ubiquitous and accessible as text. When it becomes the medium through which we understand, analyse, and share information about our world, it changes how we document knowledge and make sense of complex systems.
This analysis is based on comprehensive research conducted between December 2022 and June 2023, examining the current state of camera hardware, market developments, and emerging applications across the security, automotive and consumer electronics sectors.