One of the first things you learn in geology is the way different materials organize internally. Quartz is a crystal, with its molecules arranged in repeating lattices extending in all directions. If you break one, it’s easy to predict the angles of the fracture planes. By contrast, obsidian is amorphous. It is basically volcanic glass frozen mid-flow, with its molecules lacking any form of long-range order. Both are valid forms of matter but have different internal structures.
This isn’t really about me reminiscing the good ol’ college days or how much I can recall of my undergraduate training. I spent my childhood tinkering with computers, developing a natural intuition for how hardware and software systems work from the ground up, long before I knew what social science was. These two backgrounds unexpectedly converged while working with a massive YouTube dataset for my dissertation, leading me to (re)question something fundamental about how we think about social phenomena.
For decades, empirical social science has been chasing structures. We operationalize constructs, build measurement scales, specify regression models, and report R-squared values. The implicit assumption underlying this entire enterprise is that social phenomena possess an underlying architecture that can be uncovered, mapped, and generalized. We tend to treat social constructs like crystals that have repeating, reliable internal structures that can be understood through systematic examination. But what if we’ve been ignoring alternative specifications of social scientific constructs?
This question has been nagging at me while wrangling the YouTube dataset. Over a year of data collection yielded 80M+ observations of trending videos across all locales, amounting to over 100GB of data. The sheer scale was overwhelming, but my real concern was the efficiency of the storage format. Then, an uncomfortable realization hit me. I had become too invested in thinking about data, statistics, and social science research in terms of tables and structured formats. Despite being familiar with unstructured data and non-tabular structures, it somehow eluded me that there are fundamentally different ways to conceptualize relationships between data points. One of the projects in my pipeline is to study the spatiotemporal dynamics of content diffusion on YouTube. Traditional techniques seemed suboptimal because of the complexity, censoring, interruptions, and extensive scale. The available solutions were primarily deep learning approaches. I was initially dismissive. They didn’t fit the notion of social scientific training I had internalized, that methods and models must be “explainable” to be deemed worthy or valuable. The problem wasn’t that I lacked the right data, tools, or questions. The problem was that my representation of the data and how I posed questions were limited by a narrow, crystalline way of thinking about how social scientific constructs should be operationalized. This is an outcome of the training, where we learn to see social reality through particular methodological lenses that privilege certain kinds of structures over others.
The Crystalline Bias
How we collect and store data imposes restrictions on how we think about relationships between variables or constructs, which in turn constrains the analytical procedures we can apply. Similarly, how we conceptualize our constructs restricts how we design data collection and storage. The typical process goes something like this – theory about relationships between constructs > operationalization > research design and measures > analysis. This pipeline has served social science reasonably well, but it explicitly deems that social phenomena can be decomposed into variables with specifiable relationships. Of course, it is also reflective of the ontological foundations of most social scientific approaches, where social reality has a structure we can discover, formalize, and replicate to varying degrees.
Empirical social science attempts to emulate the rigor of the scientific method. Mathematics is rightly understood as foundational to scientific epistemology, but its assumptions impose conditions and bounds on how we describe the world. Using mathematics to study physical laws is mostly unproblematic because we have laws that meet the same requirements of rigor as axiomatic principles in mathematics. When we translate that epistemology to social reality, however, we introduce an additional layer of noisy complexity. We end up struggling to explain more than 25% of the variance of social phenomena using mathematical models. This sounds suboptimal, yet the fact that we can explain even that much is remarkable and strengthens the case for mathematical approaches to studying social phenomena.
But that still leaves 75% of the variance unexplained. Interpretivist and constructivist epistemologies have long pointed out the limitations of positivist approaches, arguing that social reality cannot be fully captured through mathematical formalization. This critique is entirely valid and well-established. The problem I’m explaining isn’t that (post)positivist epistemologies have blind spots; scholars across traditions have demonstrated this extensively. Rather, the issue is how certain method-epistemology pairings have become ossified. Interpretive methods are paired with constructivist epistemologies, quantitative methods with positivist ones, and computational approaches are treated as either atheoretical or reducible to one camp or the other. This territorial division obscures a different question. What if the limitation isn’t inherent to mathematical or computational approaches themselves, but rather to the crystalline assumptions embedded in how we’ve been applying them?
The standard explanation for unexplained variance is that there’s too much randomness in human behavior at both individual and group levels. Social scientific methods treat this randomness as yet another parameter to be controlled for. It’s framed as error, as a deviation from the norm. This stems from theoretical thinking founded on principles of physical laws, that processes should be repeatable under the same conditions. Structures should scale across individuals, groups, and societies. If reliability and repeatability are the foundations of social scientific inquiry, then tabular thinking and data representation make perfect sense. That’s been the norm. That’s how social scientists have been trained for decades. But we now have computing power that allows us to think beyond tabular structures and analyze unstructured data. This is where my geology background becomes unexpectedly relevant.
Lessons from an Inexact Science
Geology has always been called an “inexact science.” I used to think this meant it was more descriptive than other natural science disciplines, which is partly true. But there are deeper parallels between how geology approaches science and how social scientists could benefit from similar thinking. The conceptualization of the nature of materials, that is crystalline, amorphous, or somewhere in between can be particularly insightful. The same paradigm of finding theoretical structure that can be scaled informs geological research. Amorphous materials like obsidian, opal, and various glasses are treated as aberrations, receiving little attention compared to crystalline structures. They’re neglected primarily because they lack repeating, reliable structures that can inform inferences at higher levels of abstraction. But amorphous materials aren’t devoid of structure. They possess localized structures but without a guarantee that these patterns will be reflected globally. It’s a spectrum, not a binary.
This is where I wonder whether social constructs and our approaches to operationalizing them suffer from the same crystalline biases? What if we acknowledge that certain constructs tend to be more amorphous than others? They might have fairly well-defined local structures but no significant global structure. Methods like network analysis already operate somewhat on this principle. Deep learning methods give us unprecedented tools to represent and analyze these amorphous constructs in ways that tabular, crystalline thinking cannot. This might require rethinking how we pose research questions based on data representation, observation, and analysis. Perhaps the reason large language models handle textual data better is because there’s closer alignment between the construct and the method; both are fundamentally amorphous.
Stochastic modeling attempts to quantify uncertainty, which seems like it should address the amorphous nature of social phenomena. But the thinking remains largely crystalline. Randomness becomes another component that can be modeled as a probabilistic distribution. It’s still fundamentally about finding structure, just structure that includes variance parameters. Methods like multilevel models and weighted regressions account for localized dependencies to some extent, but the focus remains on global effects. If we are to draw conclusions about global effects of social constructs, we might need to combine both amorphous and crystalline construct analysis, much like in geology. A friend pursuing his PhD at a cutting-edge geology lab recently validated this intuition. The future of exploration geology lies in incorporating relational and unstructured analysis methods like network analysis and deep learning alongside established geostatistical methods. The parallel is striking.
Questions and More Questions
All of this leaves me with more questions than answers.
Can we think of social scientific constructs and theory in less crystalline terms for certain cases? How do we incorporate this into empirical thinking and combine it with existing tabular paradigms? What would be the onto-epistemological foundations of such an approach?
How do we handle explainability and transparency in amorphous methods? Is the demand for explainable AI itself an artifact of crystalline thinking biases? Deep learning techniques, particularly with textual data, operate in inherently high-dimensional spaces that are difficult to comprehend. Perhaps over time we’ll develop intuitions for these amorphous structures, making complex operations less mysterious. But until we develop better facility with unstructured data analysis, how do we address the very legitimate concerns about biases creeping into these models while not completely abandoning the idea of embracing uncertainty and amorphousness in constructs and methods?
The Way Forward
Lest it be misconstrued, I’m neither evangelizing deep learning methods nor calling for abandoning established social scientific thinking. If anything, this opens opportunities for creative methodological ideas that help us derive better insights from both tabular and reductionist data as well as unstructured data. Even amorphous substances have localized structures. The question is whether we’re willing to reconsider our foundational assumptions about what we’re studying.
For most of social sciences’ history, we’ve assumed social phenomena are fundamentally crystalline. We’ve built entire research infrastructures around this assumption like survey instruments designed for tabular analysis, theories specifying structured relationships, statistical techniques optimized for variance decomposition. But social reality might be more amorphous than we’ve made it out to be and tried to study accordingly. The unexplained variance in our models might not entirely be noise or error. It might be a signal that our crystalline methods simply cannot capture in tabular forms. Or more aptly, we are yet to develop a sophisticated enough vocabulary to translate those signals to theoretically rich insights.
The computational turn in social science offers more than just new tools. It offers an opportunity to reconsider the ontological foundations of what we study. We can continue refining crystalline methods while simultaneously developing amorphous approaches. The goal isn’t to replace one with the other but to recognize that different aspects of social reality might require different conceptual frameworks and analytical strategies. This might create some dissonance. Amorphous thinking feels less rigorous, less scientific. But that discomfort itself might be revealing. If our definition of rigor is so tightly bound to crystalline assumptions that we cannot accommodate amorphous phenomena, then perhaps we need to start thinking a bit more deeply about what the criteria for rigor in social sciences are. Ultimately, the question isn’t whether social science should become more computational or less. It’s whether we can develop the conceptual flexibility to better align our methods to the actual structure of what we’re studying. Geology taught me that not everything in nature follows a repeating lattice. Some of the most interesting materials are the ones that are frozen mid-flow.
Perhaps the same is true for social life.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.