The Problem
Databricks was founded in 2013 by the team behind Apache Spark — Ali Ghodsi, Ion Stoica, Matei Zaharia, and several others from the UC Berkeley AMPLab. The initial commercial premise was straightforward: Spark was becoming the dominant distributed-compute framework for large-scale data processing, and someone should provide a managed cloud service around it.
That framing turned out to be too narrow. Under Ghodsi's leadership after he took over as CEO in 2016, Databricks progressively expanded its ambition — from Spark-as-a-service to a full data platform combining data lakes, data warehouses, ML workflows, and (in the 2020s) generative-AI development. The architectural bet was that the historical split between the "data lake" and "data warehouse" categories was unnecessary and could be collapsed into a single "lakehouse" architecture. Databricks named the category and worked to define it.
The Journey
Ghodsi's academic background is in distributed systems. He completed his PhD in Sweden and joined the Berkeley AMPLab as a postdoctoral researcher, where he worked on systems including the resource-scheduling paper Mesos. He co-founded Databricks in 2013 alongside Ion Stoica, Matei Zaharia (Spark's original creator), and others from the lab, and took over as CEO in 2016.
Databricks' fundraising trajectory has been extensively documented in Bloomberg, The Wall Street Journal, and The Information: successive rounds from Andreessen Horowitz, NEA, Coatue, Fidelity, and others, with valuations moving from single-digit billions to $43 billion in mid-2023 to over $60 billion in 2024 to the largest AI-adjacent private valuation of the era in 2025.
Responses
Join the conversation
You need to log in to read or write responses.
No responses yet. Be the first to share your thoughts!