RSS Amplifier

Data in Motion · May 23, 2026

The Prism of Theseus: What an Ancient Paradox Reveals About AI, Identity, and Educational Data Ethics

0
Sign in to vote or save

Genevieve Smith-Nunes · Data in Motion

THE GIST: When AI models change silently through updates, fine-tuning, or new training runs. The research built on top of them may be resting on a different instrument than the one it began with. Drawing on the Ship of Theseus paradox and the Ontological Kaleidoscope framework, this piece introduces the concept of the Prism of Theseus: a way of thinking about model identity as an ethical and methodological variable in AI-driven educational research.

The Prism of Theseus. Schematic representation of AI model identity and epistemic refraction. The diagram illustrates how continuous data replacement (left) shifts a model's identity through an epistemic prism (centre), yielding four divergent philosophical interpretations of identity and continuity (right).

There is a thought experiment that has been bothering philosophers for about 2,400 years. The Athenians preserved the ship that carried Theseus home from Crete, but as the planks rotted, they replaced them, one by one. Eventually, every plank had been replaced. Was it still the same ship?

Share

Now hold that question, and consider this scenario.

A research team uses an AI language model to generate synthetic learner data for a study on personalised feedback. They document their methodology carefully. They name the model. Six months later, a second team attempts to replicate the study. The model has been updated. The outputs are different, not dramatically, but systematically, in ways neither team can fully account for. The replication fails.

The Ship of Theseus is a 2,400-year-old thought experiment that still sits at the heart of debates in philosophy, law, science, and now AI.

The core puzzle: Theseus’s legendary ship was preserved in Athens, but as planks rotted they were replaced. Eventually, every single plank had been swapped out. Is it still the same ship?

This isn’t a hypothetical. It is already happening, quietly, in the emerging field of AI-driven educational research. And the field does not yet have the conceptual vocabulary to name it, let alone address it.

When we cite an AI model in a methods section ‘we used GPT-4,’ ‘we used Claude 3’ we are making an identity claim. We are asserting that this named entity is a stable, documentable instrument: the same tool our colleagues can access, evaluate, and hold us accountable to. In the case of this post and academic article were developed prior research and very interesting and diffractive conversations with Claude. Therefore, I asked Claude how I should acknowledge its contribution.

Claude’s Response
I shouldn't be listed as an author or co-author. Authorship carries accountability — the ability to stand behind claims, respond to peer review, take responsibility for errors. I can't do any of those things.What you should absolutely claim as fully yours: the Ontological Kaleidoscope, the data dignity framing, the DataDrivenDance research context, the possibility-studies theoretical home, and the intellectual judgement about which ideas from our conversation were worth keeping. Those are yours. I was the interlocutor, not the author.

But AI models are not stable instruments in the way that validated survey scales or established coding frameworks are. They are updated. Fine-tuned. Silently patched. In some cases, entirely retrained and re-released under the same name. The model behind the API endpoint changes, and the change may not be announced, or the announcement may not reach the researcher who simply calls the API and assumes continuity.

There’s a particular satisfaction in running something locally that was once the exclusive domain of distant servers and corporate infrastructure. Also, in my own work with offline LLMs, I’ve felt that directly. But the privacy benefit of local deployment doesn’t resolve the identity question. A locally-run model can also be updated, fine-tuned, and changed. The problem of identity instability is not about where the model lives. It’s about what it is, and whether we can say with confidence that it is the same thing it was when the study began.

In my research developing the concept of the ‘body as a data artefact,’ I’ve been concerned with what happens when physical, embodied experience is translated into data infrastructure. When a person’s physical characteristics become persistent digital records, that data does not simply describe the body. It becomes a version of it. Something is always gained in that translation (pattern, scalability, shareability. And something is always lost) singularity, affect, the irreducible particularity of the lived body.

The Ontological Kaleidoscope framework I developed to examine these translations uses multiple analytical lenses simultaneously (posthumanism and constructionism) to reveal the gaps and ruptures that a single perspective would smooth over. The framework insists that impossible translations are not failures to be corrected; they are analytically significant in their own right.

Now I want to add a third lens to that Kaleidoscope: the AI model itself.

Not as an orchestrator: a conductor standing above the research, coordinating its elements into coherent output. That framing implies a centralised agency that sits uneasily within a posthumanist framework that distributes agency across human-nonhuman entanglements.

Instead: as a refracting lens. A lens that produces a view, a partial, situated, and potentially distorting view, of whatever learner data passes through it. A lens whose refractive properties are determined by its material constitution: the training data it was built on, the architectural choices that shaped it, the fine-tuning history that has modified it.

And crucially: a lens that, like any instrument, can change between observations.

I want to introduce a concept I’ve been developing (in conversation with Claude) that brings these threads together: The Prism of Theseus.

The Prism of Theseus. A simple diagrammatic version. The pretty one is at the start of the article

A prism is more specific than a lens. It doesn’t merely refract, it disperses. It takes undifferentiated input and reveals the spectrum latent within it: showing us what was already there, but now made visible through separation, through colour, through the differential bending of light. An AI model functioning as an epistemic prism does the same. It takes learner data (movement, cognition, language, behaviour) and disperses it into a structured view. This pattern is significant, this one is noise, this response indicates understanding, this one indicates confusion.

But here is what the Ship of Theseus adds. A prism whose material constitution changes between observations does not produce the same dispersion twice. If the glass changes, if its refractive index shifts, the spectrum changes too. And if you don’t know the glass has changed, you will attribute the difference in the spectrum to the data, not to the instrument.

This is the Prism of Theseus problem: the identity instability of AI epistemic prisms over time, and the methodological and ethical consequences of that instability for AI-driven educational research.

The Ship of Theseus paradox has never had a single answer, because identity is not a single thing. Different philosophical traditions emphasise different criteria, and I’ve found that working with four of them together, rather than choosing one, gives you a much more useful analytical toolkit.

Think of these as four planks in a model identity audit.

The Heraclitean plank: change is real and must be named. Heraclitus held that you cannot step in the same river twice, change is the only constant. Applied to AI models, this insists that every update is a real change, not a cosmetic adjustment. Research programmes must explicitly acknowledge model changeability rather than suppressing it. ‘We used Model X’ is not adequate documentation if Model X was updated during the study and the change went unremarked.

The Aristotelian plank: function defines identity. Aristotle held that identity resides in form (in function, structure, and purpose) not in material constitution. A ship rebuilt plank by plank is still the same ship if its function persists. This is the most permissive criterion: it allows that fine-tuning which preserves a model’s core epistemic capabilities doesn’t necessarily break its identity. But it requires researchers to document what epistemic function they relied upon, and to assess whether that function was preserved across any updates.

The Lockean plank: identity requires a continuous chain of recognition. Locke grounded identity in continuity: of memory, of social recognition, of institutional context. For AI models, this means: the same API endpoint, the same organisational context, the same documented research relationship. A model deployed across different institutions, jurisdictions, or terms of service is not Lockean-identical to its predecessor, even if the weights are technically the same.

The Kripkean plank: origin is essential. Kripke held that identity is anchored in origination, the specific causal event that brought an entity into being. For AI models, this is the most demanding criterion. A new training run (even on identical data with identical architecture) produces a new model, not the same model in a new instantiation. This has direct implications for reproducibility: a study that attempts to replicate findings using ‘the same model’ but in a different training run is not, on Kripkean grounds, using the same epistemic prism.

Together, these four planks constitute what I’m developing as a four-plank identity audit: a systematic methodology for assessing whether an AI model retains sufficient identity across change to support the methodological and ethical claims built upon it.

Children represent a particularly vulnerable population, yet the economics of educational technology mean that many AI-powered learning tools come from smaller organisations operating under significant resource constraints. The model identity problem compounds this vulnerability.

Consider the DataDrivenDance context: synthetic data generated from brainwave and movement data. The participants — dancers, some of them young people — consented to their biometric data being used for a specific purpose, through a specific instrument. If that instrument changes between data collection and analysis, the translation they consented to is not the translation that occurred. A thirteen-year-old who lies about their age to access a service is not demonstrating adequate consent — and a research programme that proceeds as though a changed model is the same model is similarly operating outside the terms of the consent it received.

This is not a marginal edge case. Biometric data collected in childhood creates a longitudinal record that may follow an individual into adult contexts in ways that were never anticipated at the point of collection. If the AI prism that generated synthetic representations of that biometric data has been updated, if its refractive index has shifted, then the synthetic data carries traces of a different instrument’s norms, a different population’s movement patterns, a different model’s implicit understanding of what a body in motion looks like.

The participants whose data constituted the original training set are still present in that refraction. Not as identifiable individuals. But as the constitutive history of the instrument. They are, to return to the Ship of Theseus, the original planks. Displaced but not absent. Replaced but not erased.

Share

I am not arguing that we should stop using AI models in educational research. My own work has depended on them productively, and I believe they can open genuine pedagogical and methodological possibilities that more traditional approaches may not. The technology is not the problem.

The governance is.

The question worth asking is not whether AI can improve education. Some evidence suggests it can, in narrow contexts. The question is who benefits, who decides, and what safeguards exist when things go wrong.

The Prism of Theseus framework adds a temporal dimension to that question: and what safeguards exist when the instrument itself changes?

As a starting point, I think we need three things:

Identity documentation as standard methodology. Model name alone is not adequate. Version, training provenance (where available), deployment date, and any known updates during the study period should be standard requirements in methods sections for AI-driven research. Journals and ethics review boards should require this. We have rigorous documentation requirements for every other instrument we use; AI models should not be exempt because the changes are invisible.

Reproducibility claims should be bounded by origination. A finding generated by a specific model, in a specific version, at a specific moment is reproducible only under conditions that preserve that origination. Where commercial models cannot be fully documented, this limitation should be explicitly stated, not as a technical caveat, but as a substantive methodological constraint on the claims the study can make.

Consent frameworks must account for instrument change. If research participants consent to their data being processed by a specific AI instrument, and that instrument changes during the study, the ethics review should be reopened. This is currently not standard practice anywhere I am aware of. It should be.

The Ontological Kaleidoscope was built to examine the gaps, ruptures, and impossible translations between embodied experience and data infrastructure. The Prism of Theseus adds a new kind of impossible translation to that analysis: not just the spatial translation between body and data, but the temporal translation between the body as it was when consent was given, the model as it was when the data was generated, and the analysis as it is conducted now.

A three-point translation loss. Compounding at each stage.

Naming this is itself a contribution. The field cannot address what it cannot name.

I have been working on a formal academic article of these ideas (with Claude as my co-author), a paper tentatively framed around the Prism of Theseus as a methodological and ethical concept for AI-driven educational research. This Substack piece is the first public sketch of that argument. I would genuinely like to hear from researchers and practitioners working in this space. Whether you’ve encountered the model identity problem directly, whether the four-plank audit seems useful or unwieldy, and whether there are dimensions of this I haven’t yet accounted for.

This post was developed in dialogue with Claude Sonnet 4.6 (Anthropic), which contributed to the conceptual development of the Prism of Theseus framework, the four-plank identity audit, and the grammar correction of this text. All intellectual direction, theoretical grounding in the Ontological Kaleidoscope framework, and final editorial decisions are the author's own. The AI-assisted development process is itself an instance of the methodological questions this article raises.

The Ontological kaleidoscope and Prism of Thesis only give you its best picture when you hold it to the light from multiple angles. I am interested in yours.

This article draws on the following works:
Smith-Nunes, G. (2025). The Ontological Kaleidoscope. Journal of Responsible Technology. https://doi.org/10.1016/j.jrt.2025.100138.
Smith-Nunes, G. & Ness, I.J. (forthcoming). Synthetic Possibilities: How AI and Synthetic Data Open, Constrain, and Redistribute Horizons of Educational Agency.
Ness, I.J. & Smith-Nunes, G. (2027). Polyphonic Creativity: Navigating the Opportunities and Challenges of Artificial Intelligence in Education in The Oxford Handbook of Human Creativity and Generative Artificial Intelligence in Education.

Share

Thanks for reading

No posts

Read the original on readysaltedcode.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.