Over the past eighteen months, DataSF has undergone a fundamental transformation. Our mandate has expanded from merely publishing the City’s data to building and operating the essential infrastructure that enables the City to utilize it effectively. This evolution required significant institutional work, such as defining the team’s authority and scoping its remit across departments to create the conditions for a robust platform to exist. As San Francisco’s Chief Data Officer, I am excited to introduce our new direction, the functions we are building, and how we intend to engage with our community.
Our progress is built upon a strong foundation. DataSF was established on the principle that city data belongs to the public and that open publication creates compounding civic value. This vision led to the creation of our open data portal, comprehensive publishing processes, and a citywide framework that treats data as a strategic asset. Today, hundreds of datasets are available across dozens of departments, supported by millions of monthly API calls. While this foundation remains vital, the city’s growing needs now demand a more sophisticated approach to data management.
Modern challenges require departments to share information across various teams and departments to coordinate services, allocate resources, and evaluate program effectiveness. Consequently, the systems we build today require higher standards for governance, security, and ethical use. Because decisions are increasingly made through direct system queries rather than static reports, we have moved beyond the traditional “publish the dataset” model to a structure organized around four interdependent functions.
Platform Engineering builds and operates the Unified Data Platform (UDP), the citywide infrastructure that centralizes data under common quality and security standards. By providing this foundation, we eliminate the need for one-off integrations and allow the City to address problems with necessary speed. We are also looking for a fantastic builder to join this team!
Enablement focuses on analytics engineering, transforming raw operational data into modeled, trusted data products. This crucial layer ensures that data is not just technically available but practically useful for departments to build upon.
Data Science tackles complex analytical problems through predictive modeling, entity resolution, and rigorous evaluation methods. Now formally integrated into the platform, this function empowers the City to make better decisions using existing data collections.
Operations serves as the institutional plumbing, managing governance, policy, and data-sharing agreements. This work ensures that our citywide data program remains reliable and durable across various administrations and budget cycles.
These functions are deliberately linked: a platform without enablement produces unused data, while data science without a platform lacks compounding impact. By working together, these teams create a sustainable ecosystem for public-sector data. To continue this mission, we are actively hiring professionals that are infrastructure builders, modelers, and analysts who want to work on systems that directly affect how San Francisco delivers services.
Finally, we view this blog as a two-way channel for public engagement. We want to hear your perspectives on data problems worth solving or datasets that could change your understanding of the city. Moving forward, this publication will be the space where we transparently share what we are building and learning. We are committed to an honest dialogue about the realities of modern data infrastructure as we work to serve the people of San Francisco.
Soumya Kalra
Chief Data Officer, City and County of San Francisco
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.