PhilPapers: Recent additions to PhilArchive
Alhosseini Almodarresieh, Seyed Alireza: The Ship of Theseus in Large Language Models: A Framework for Quantifying Model Identity and Alignment Continuity under Continual Fine-Tuning
0Sign in to vote or save
This page did not load. You can still read it on the original site — the toolbar below keeps your place in the directory.
Continual fine-tuning changes an LLM’s weights directly, unlike ordinary software upgrades. Drawing on the “Ship of Theseus” paradox, we ask when a fine-tuned model ceases to be usefully “the same” model with respect to its original alignment properties. We define Model Identity Continuity and a candidate index, the Theseus Stability Metric (TSM), combining a normalized, per-layer…

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.