Synthetic Data Curation Architecture for Agile SLM Adaptation to Emerging Programming Language Specifications
Adapting a 1.5B-parameter code model to Ruby 3.4 for approximately €4.20 in training and evaluation compute.
| Catalog ID | MDSW 1 |
| Category | Applied AI Research |
| Project Status | Maintained (the pipeline is still in active use) |
| Involvement | 2026 · Sole researcher |
| Stack | Python, Terraform, Google Cloud Platform |
| Models & Datasets | Training dataset · LoRA adapters · Eval logs |
| Paper | Under review, Automated Software Engineering (Springer) |
Large language models know Ruby the way a widely read person knows a language they haven’t spoken in a while: broadly, but out of date. Ruby 3.4 did not exist when they were trained, and no amount of general fluency writes syntax a model has never seen. One answer is to wait for the next large model to catch up. This thesis asked a narrower, cheaper question: could a small model be taught the missing pages directly, fast and cheaply enough that a team could do it as routine work?
Answering that meant building a pipeline that could run the experiment over and over without getting expensive: synthetic data generation, a seven-stage quality filter, LoRA fine-tuning and automated evaluation. It ran across 8 model architectures from 1.5B to 9B parameters and produced 96 fine-tuned models, 156 evaluation runs and almost 33,000 logged evaluation records. A control group, trained on the same number of unfiltered examples as the curated set, separated what curation bought from the effect of simply seeing fewer, cleaner tokens. A separate diagnostic, the IC/EC split, asked why a model failed: because it did not know the answer, or because it could not hand the answer over in the right form.
The clearest result: a fine-tuned 1.5B model, small enough to run on a phone, closed most of the gap to a generalist model 4.7 times its size, for an amortized cost of about €4.20 in training and evaluation compute per specialist model. Curation did not win everywhere. An uncurated sample of the same size solved clearly more tasks outright, but curation pushed idiomatic style to 8.3, above the teacher model’s own 7.89. The IC/EC split showed that base models already had much of the reasoning and lost their points on formatting: repairing only the form of their answers raised execution pass rates by 21.1 percentage points on average. LoRA fine-tuning is what taught the models to get the form right on their own. The pipeline gives a repeatable way to build a narrow expert model in-house, without depending on an external API.
Findings from this work are under review at Automated Software Engineering (Springer). Originated as an MSc thesis in Computer Engineering, Istanbul Bilgi University, supervised by Assoc. Prof. Dr. Savaş Yıldırım.
Record: YÖK National Thesis Center, Tez No 1024515. YÖK records have no direct link; search by the number.