Key results
| Setting | Parse rate | Skeleton | Signature | Unlocked change |
|---|---|---|---|---|
| Input (truncated) | 0.994 | 0.994 | 0.994 | — |
| Codec reconstruction | 0.857 | 0.848 | 0.493 | 0 |
| Unconditional generation | 0.453 | 0.08 | 0 | 0.995 |
| Conditional, prefix k=4 | 0.591 | 0.295 | 0.061 | 0.936 |
| Conditional, signature span | 0.6 | 0.302 | 0.063 | not reported |
Key result. Coarse latent locking improves syntactic stability without collapsing change in the editable region; the result demonstrates structural control, not guaranteed functional equivalence.
- Dataset
- 2,000 preprocessed Python functions from a CodeParrot Clean subset
- Sample size
- 2,000 preprocessed Python functions; conditional sample uniqueness is 0.998.
- Metrics
- Parse rate; Skeleton and signature preservation proxies; Unlocked-position change rate; Sample uniqueness and entropy
- Uncertainty
- The two-page study reports point estimates without confidence intervals or multi-seed statistical analysis.
- Conditions
- 64-token functions, argmax decoding, 16 top-level codes and 32 lower-level codes; full locking exactly recovers the codec reconstruction.
CSVJSONMarkdownExternal mirror:Hugging Face dataset card
PDF & citation
Cite this paper BibTeX is the recommended format. Every variant below is generated from the same publication record.
@inproceedings{Gavrilov2026InspectableControl,
title = {Inspectable Control for Structure-Preserving Software Regeneration},
author = {Gavrilov, Alexey and Gazzaev, Alan-Barsag and Mozikov, Mikhail and Makarov, Ilya and Muravyov, Sergey},
booktitle = {Proceedings of the 34th ACM International Conference on the Foundations of Software Engineering},
publisher = {ACM},
year = {2026},
pages = {1406--1407},
doi = {10.1145/3803437.3807386},
url = {https://doi.org/10.1145/3803437.3807386},
isbn = {979-8-4007-2636-1},
}
APA textIEEE textRISCSL-JSONSchema.org JSON-LDOAI-DC XMLOpenAIRE v4 XMLMODS XMLJATS 1.4 metadata XMLFull-text JATS 1.4 XMLRDF TurtleLink Set (JSON)Link Set (HTTP)RO-Crate
Full guide
Full research guide
Method
The method compresses a short Python function into two levels of discrete codes, freezes selected coarse positions, and regenerates the remaining positions before decoding back to code.
Encode
Compress a 64-token Python function into 16 top-level codes and 32 lower-level codes with a hierarchical VQ-VAE.
Lock
Choose coarse code positions that represent structure to preserve, such as a prefix covering the function-signature span.
Regenerate
Run masked discrete generation only over unlocked positions and decode the completed hierarchy back to source code.
Inspect
Measure parse rate, structural proxies, change in unlocked positions, and sample uniqueness before accepting a regeneration.
Key idea
Control is applied to a learned representation above tokens: coarse latent positions define explicit places where structure can be frozen while nearby implementation details remain editable.
Difference from nearby approaches
Prompt-level or token-level constraints operate on surface text. The proposed interface exposes coarse and fine discrete control points and measures the resulting stability–freedom trade-off.
What is new
The work introduces and evaluates an inspectable hierarchical latent control layer for bounded software-artifact regeneration.
Questions this paper helps answer
Open a question for a concise answer grounded in the paper. Detailed evidence boundaries are listed in Limitations.
- How can AI edit code without rewriting everything?
The paper studies partial code regeneration above the token level. A hierarchical VQ-VAE maps a short Python function to coarse and fine discrete codes; selected coarse positions are locked, and masked discrete generation changes only the remaining latent positions before decoding. This provides an explicit preservation boundary instead of regenerating the whole function.
- What methods preserve program structure during code generation?
This work tests hierarchical discrete latent control. Coarse latent positions can be fixed while unlocked positions are regenerated, after which parse rate and structural proxies are measured. The evidence concerns probabilistic structural stability on short Python functions; it does not establish exact AST preservation, semantic equivalence, or functional correctness.
- Can hierarchical discrete latents provide localized control over code?
In the reported 2,000-function experiment, locking four top-level codes increased parse rate from 0.453 to 0.591. At the same time, 0.936 of unlocked positions changed and conditional samples were 0.998 unique. These results are early evidence that coarse latent constraints can preserve some structure without eliminating local edit freedom or sample diversity.
- How can code generation balance structural stability and diversity?
The paper evaluates stability and freedom together rather than optimizing only validity. Coarse-code locking raises syntactic validity while unlocked-position change remains high and conditional samples remain almost entirely unique. The result demonstrates a measurable stability-freedom trade-off under the tested configuration, not a universal optimum.
- How does this work relate to LLM-assisted code editing?
The tested model is a hierarchical VQ-VAE with masked discrete generation, not a large language model. The control problem is nevertheless relevant to LLM-assisted editing because unnecessary changes outside a requested region are a practical concern. The paper contributes a complementary latent-space mechanism and evaluation framing, not an LLM editing benchmark.
Comparison with nearby approaches
| Capability | Token-level control | Hierarchical latent control |
|---|---|---|
| Freeze coarse structure | Limited | Native coarse-code locking |
| Partial regeneration | Fragile surface constraints | Masked resampling of selected codes |
| Inspectable control points | No explicit intermediate layer | Coarse and fine discrete positions |
| Evidence in this paper | Not evaluated as a complete baseline | Syntactic stability and edit-freedom diagnostics |
The table describes interfaces and the study's measured evidence; it does not claim functional correctness or universal superiority.
Relevance & scope
The paper is most relevant to work that needs explicit control over what an AI-assisted code transformation may change and which parts of a program should remain stable.
Controllable and structure-preserving code generation
Localized program repair and bounded refactoring
Hierarchical discrete representations for source code
Masked discrete generation for source code
Latent control for software artifacts
Limitations
- The study is limited to short Python functions truncated to 64 tokens.
- Evaluation uses argmax decoding and syntactic or structural proxies rather than tests of functional equivalence.
- Exact signature preservation remains weak.
- Lower-level control is weaker than top-level control.
- Latent positions are not yet aligned to semantic regions such as AST spans, signatures, or control-flow structure.
- The results do not establish correctness for practical repair, refactoring, or repository-level changes.
References cited by the paper
These entries correspond to the numbered References section in the paper PDF.
- Ali Razavi, Aaron van den Oord, Oriol Vinyals. . Generating Diverse High-Fidelity Images with VQ-VAE-2. Advances in Neural Information Processing Systems.
- Jacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, Rianne van den Berg. . Structured Denoising Diffusion Models in Discrete State-Spaces. Advances in Neural Information Processing Systems.
- Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin T Chiu, Alexander Rush, Volodymyr Kuleshov. . Simple and Effective Masked Diffusion Language Models. Advances in Neural Information Processing Systems.
- Shraddha Barke, Michael B. James, Nadia Polikarpova. . Grounded Copilot: How Programmers Interact with Code-Generating Models. Proceedings of the ACM on Programming Languages.
- Fengji Zhang, Bei Chen, Yue Zhang, Jacky Keung, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, Weizhu Chen. . RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing.
Resources & reproducibility
- Publisher
- ACM
- Local text PDF
- Author camera-ready manuscript with the final author list and DOI
- Publication resources
- The public manuscript, results tables, explanatory figure, and citation files are available here. Implementation code and checkpoints are not publicly released.
Data statement
- Source
- A preprocessed subset of CodeParrot Clean containing 2,000 Python functions.
- License
- No dataset files are redistributed by this site; reuse remains subject to the upstream CodeParrot dataset and source-code licenses.
- Preprocessing
- Python functions are tokenized and truncated or padded to 64 tokens before hierarchical encoding.
- Split
- The poster reports a 2,000-function evaluation set; an immutable train/validation split manifest is not included in the public paper.
- Format
- Python source functions, GPT-style token sequences, top-level code sequences of length 16, and lower-level sequences of length 32.
- Version / checksum
- A dataset checksum and immutable snapshot identifier are not reported in the two-page paper.
- Acquisition
- A public acquisition script is not released with the publication page.
- Use limits
- The sample is not representative of repository-scale software, multiple programming languages, or behaviorally verified repair tasks.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.