Beneath the internet, there has been an unwritten agreement for years. Apps, platforms, and services are available to you for free. They obtain your information, including your location history, late-night searches, shopping preferences, and health concerns. There was never a written record of the arrangement. The majority of people didn’t give it much thought. And somewhere in that quiet, a multibillion-dollar industry was subtly constructed on top of the digital lives of regular people.
There is beginning to be significant opposition to that quiet arrangement. A small but increasing number of startups are now creating what they refer to as “personal data vaults,” which are user-controlled, encrypted storage systems that allow users to control who can access their data and, increasingly, receive payment when businesses request access. The basic argument is almost unbelievably straightforward: if your data is valuable, why aren’t you making money off of it?
One of the more closely watched companies in this market, Caden, allows users to extract behavioral data from services they already use, like Uber, Netflix, and Amazon, and then decide whether to share particular segments with advertisers. In 2023, the business raised $15 million in a Series A round with the support of Yahoo co-founder Jerry Yang.
John Roa, the platform’s founder, created it after observing something that didn’t seem right: for about 25 years, users had been providing massive amounts of personal data in exchange for “free” services, and the businesses that received it had amassed that value without being very transparent about what they were doing with it. In hindsight, Roa’s argument is nearly obvious, so it’s not radical. There was always value in the data. Who got to make that decision was the question.
Throughout the space, the model is different. Some startups compile anonymized data from thousands of users and market it as a package. The reasoning behind this is that while individual data points aren’t particularly valuable, patterns found among 100,000 users are.

Others are experimenting with what they refer to as royalty meters, in which users receive a small payment each time a feature that a business is actively paying for benefits from their anonymized signals. Some are proposing something more akin to a data-rental model, in which users share information for a predetermined amount of time and for a specific purpose before withdrawing it. It seems as though no one has yet discovered the ideal structure. The experiments feel truly exploratory because the category is still in its infancy.
It’s important to be explicit about the scale at play. The earnings that are currently being reported are modest. Caden has disbursed “hundreds of thousands” of dollars to its user base, which seems significant until you consider that it’s dispersed among a potentially sizable number of contributors. This is unlikely to take the place of a paycheck for the majority of people. Not yet, anyway. It’s still genuinely unclear if this will change as these platforms expand and the need for high-quality behavioral data increases.
The timing is also intriguing. California, Colorado, the EU’s GDPR, and other laws on both sides of the Atlantic have pushed businesses to give customers more control over their personal data, but the majority of these laws are based on the right to opt out rather than the right to earn. These startups are now taking advantage of a kind of opening created by the regulatory environment. They are not holding out for Washington to develop a framework for compensation. They are constructing one themselves, which, depending on how much you believe the incentives involved, can be either entrepreneurially astute or a little unsettling.
Although it’s more difficult to assess, a more speculative argument is also receiving some attention. Personal archives may become truly valuable training material as AI systems become increasingly hungry for real, human-generated data and as the public internet’s supply of high-quality text starts to diminish. creative work, message histories, and voice journals. Some people are beginning to consider the potential value of a carefully curated digital archive of their own life to an AI model in the future. It’s difficult to tell if that’s wishful thinking or foresight.
It is evident that the long-held belief that sharing your data is just a cost of using the internet is under scrutiny more than it has ever been. It remains to be seen if these startups can truly fulfill their promise to turn that data into revenue on a significant scale. However, the fact that serious investors are paying attention indicates that someone thinks the math might eventually make sense.
