Cloudant
Three MIT physicists were moving petabytes of data from the world's largest physics experiment and couldn't find a single tool that worked at that scale. So they built one — and accidentally created the cloud database infrastructure that IBM eventually paid to own.
Michael Miller · 17 min read
Cloudant, YC Founder Story
Company: Cloudant Founders: Michael (Mike) Miller (Chief Scientist), Adam Kocoloski (CTO), Alan Hoffman (Director of Product) YC Batch: Summer 2008 (S08) Industry: Database-as-a-Service / Developer Infrastructure / Cloud Founded: 2008, Cambridge, Massachusetts Acquired: IBM, announced February 24, 2014, closed March 4, 2014 Total Funding: $15.1 Million (pre-acquisition) Post-Acquisition: Now IBM Cloud Data Services
The One-Line Summary
Three MIT physicists were moving petabytes of data from the world's largest physics experiment and couldn't find a single tool that worked at that scale. So they built one, and accidentally created the cloud database infrastructure that IBM eventually paid to own.
Lens 1, The Before State
Who Were They Before YC?
Michael Miller, Adam Kocoloski, and Alan Hoffman were not startup founders. They were researchers. Physicists. People whose daily work involved understanding the fundamental building blocks of matter, not building companies.
Mike Miller described himself plainly: "I'm a philosopher and a physicist by training." He was an Assistant Professor of Particle Physics at the University of Washington. Adam Kocoloski was a researcher and would become Cloudant's CTO. Alan Hoffman, who would become Director of Product, later said the early days involved "leading dual lives: the life of science, and the life of a tech startup, which I wouldn't encourage anyone to do."
Their world was the Large Hadron Collider (CERN) and the Relativistic Heavy Ion Collider (RHIC) at Brookhaven National Laboratory, two of the most complex scientific instruments ever built by humanity. The LHC generates more data in a single second than most companies process in a year. Moving, storing, analysing, and distributing multi-petabyte datasets was not an abstract challenge for these three, it was Tuesday.
The Personal Pain They Were Living
The problem was deceptively unglamorous: the tools for managing big data were awful.
Every time the trio needed to analyse a new dataset from the LHC or RHIC experiments, they faced the same friction. Data was siloed. Moving it required custom scripts. Analysing it in different locations meant rebuilding infrastructure from scratch each time. The databases available in 2007, 2008 were either too rigid (relational SQL), too fragile (didn't scale under heavy load), or simply impossible to distribute globally without heroic engineering effort.
Miller put it simply: "Scientists spend a lot of time building the technology needed to do science", only getting to use those tools for actual analysis much later. The overhead was consuming their research lives. So they started building something better, first for themselves.
They discovered CouchDB, an open-source, document-oriented NoSQL database built for distributed systems. It was promising but needed to be pushed much further. So they built BigCouch: their own open-source, fault-tolerant, clustered, globally scalable extension of CouchDB that could handle the demands of particle physics research.
They didn't plan to turn it into a company. They were just trying to do their jobs.
Why They Almost Didn't Look Like Startup Founders
Three physicists from MIT applying to a Silicon Valley startup accelerator in 2008 was genuinely unusual. YC's typical profile at the time skewed toward young hackers building consumer web apps, not tenured researchers building distributed database systems. Miller, Kocoloski, and Hoffman were academics first. They had no prior startup experience, no business background, and were still employed full-time in research when they applied.
There's a particular kind of imposter syndrome that comes with being deeply technical but having no startup credentials, no prior exit, no famous founder-friends, no network in Silicon Valley. These three had none of that. What they had was a working system handling some of the most demanding data workloads on earth, and the insight that if particle physicists needed it, so would everyone building mobile and web apps.
Key Insight for Aspiring Founders
The best startup ideas come from the most unsexy places. Nobody reads physics research papers looking for startup inspiration. But three researchers staring at petabytes of collider data saw a universal developer problem years before anyone in Silicon Valley was talking about it. Your deep domain boredom is someone else's billion-dollar pain point.
Lens 2, The Idea Origin
How the Idea Was Actually Born
There was no single "aha moment." The idea emerged through accumulated frustration over years of working with data that modern tools couldn't handle.
In 2007, while working on data problems at the LHC, which generates extraordinary volumes of collision data per second, Miller, Kocoloski, and Hoffman found CouchDB. It was elegant in design but not remotely ready for production-scale distributed workloads. So they began extending it.
What became BigCouch, and eventually Cloudant, started as internal research infrastructure. The three were solving a problem in front of them: how do we store, replicate, search, and analyse massive datasets across geographically distributed systems without everything breaking?
The pivot from "research tool" to "startup" happened when they asked a different question: "If we need this so badly, who else does?" The answer, they realised, wasn't other physicists, it was every developer building web and mobile apps that needed data to live reliably in the cloud. The specific insight was that mobile computing was about to explode, and most developers had no good way to manage the data behind their apps at scale.
Their idea in one sentence: take the infrastructure we built to survive particle physics, clean it up, and sell it as a service to anyone who needs their database to never go down.
The First Ugly Version
The first version was, strictly speaking, a research tool, not a product. BigCouch was open source, written for scientists who understood distributed systems. It had no onboarding, no documentation designed for non-physicists, no support infrastructure, and no pricing.
Getting from "research tool used by physicists at CERN" to "product a developer could sign up for in three steps" required them to completely rethink who they were building for. The original users were their colleagues. The new users were developers who had never heard of the Large Hadron Collider and didn't care.
The real insight buried in this ugly first version: they already had proof that the core technology worked at the absolute extreme end of scale. Most database startups spend years proving their system doesn't collapse under load. Cloudant's system had already survived CERN.
The Signal That Validated It
The signal wasn't a sudden surge of users. It was quieter: YC's interest itself.
When the team's work on big data caught the attention of Y Combinator in early 2008, it validated something important, the problem they were solving in academic research was commercially relevant. YC's investment in them was a signal that people outside of particle physics saw the same infrastructure gap they did.
The second signal was how developers responded when Cloudant eventually launched a free tier in August 2010. Developers didn't need to be educated on why reliable, scalable, globally distributed data was valuable. They had been struggling with exactly this problem and simply hadn't had a clean managed solution. Word spread quickly through developer communities.
The Pattern This Follows
Cloudant is a textbook example of what YC calls "research-to-startup", a pattern where deep technical work done in academia or advanced research environments produces tooling so good it has obvious commercial application. Other famous examples from this pattern include Google (search infrastructure from Stanford), and many open-source database companies. The key is that the research bona fides aren't just a story, they're a real moat. Competitors couldn't replicate the depth overnight.
You've read your 5 free stories this month.
The next YC application you write could be the one that gets in. Don't stop learning from the founders who already did it.
- ✓ 1000+ deeply-researched YC founder stories
- ✓ Unfiltered Lens breakdowns: what worked, what failed
- ✓ 2 new founder deep-dives every week
- ✓ Full Q&A library + application teardowns
Free YC databases
Everything behind this page is in our open databases
This story is one slice of the data we keep open — batch lists, rejection case studies and launch playbooks, all free to read.
Database
YC Rejection Database →
Airbnb, Stripe, Reddit — founders who got rejected first, with every source linked.
Index
List of YC Companies →
Searchable index of YC alumni by batch, from Airbnb (W09) to the newest AI startups.
Database
YC Solo Founder Database →
20+ verified solo founders, their pre-YC traction and the exact application framing they used.
Browse all free YC databases → · or start on the YCInsight homepage
Go deeper on what Cloudant did
Questions this story raises
- YC RFS Developer Tools — Specific Opportunities YC Wants Founders to Build
- YC Application 2026: The Complete Checklist (20 Steps Before You Submit)
- Complete List of YC S24 Companies — Summer 2024 Batch
- YC W19 Batch — Where Are These Companies Now
- How to Answer "What Is Your Company?" on the YC Application
- YC Application for Pre-Revenue Startups — What to Focus On
Free tools for this stage
Start from the YCInsight homepage for every YC database, founder story and Q&A in one place.
Related founder stories
Keep reading — more YC founders from Summer 2008 (S08) and beyond.
BackType
Two Canadian engineers from Toronto — who had already built and failed with one startup — walked into YC Summer 2008 with a simple idea about blog comments, and walked out having quietly invented the infrastructure that powers real-time social data analytics at Twitter scale.
Poll Everywhere
Three Deloitte consultants were so tired of watching audiences zone out during their own presentations that they built a text-message voting tool to keep people awake — then discovered they'd accidentally created the product that would kill a $15,000-per-unit hardware industry.
Posterous
Three Stanford engineers built the world's simplest blogging platform — post anything just by sending an email — grew it to 15 million users, then watched it unravel the moment they stopped trusting simplicity. One founder went on to run Y Combinator. The story of Posterous is equal parts inspiration and cautionary tale.
Vidyard
Two engineering students in Canada — one designing toilets, one sitting idle at BlackBerry — used a day-trading windfall to buy a house, build a video company nobody asked for, and accidentally discovered the real product while trying to solve their clients' confusion. Then they drove 1,200 miles with a car covered in stickers to crash a conference and corner a keynote speaker.
Treehouse
Ryan Carson watched developers graduate from college without knowing how to code for real jobs — so he spent 6 years building a blog audience first, then launched a product to that audience and hit $1.7M revenue in under 12 months, teaching over a million people to code before anyone in Silicon Valley thought online education could work.
inDinero
A 19-year-old CS student who had never held a real job, knew nothing about accounting, and dropped out of high school at 15 — built the "Mint.com for business," nearly destroyed it through arrogance, then pivoted her way to 2,686% revenue growth and the cover of Inc. magazine.
More stories every week.
Get 2 free stories in your inbox each week.