A local AI experiment
Training a Small Model to Decode Unicode Braille
Using LoRA on a single consumer GPU, character accuracy rose from 95.39% to 99.89% on the same 100 short test sequences, without an alphabet reference. Longer everyday sentences remained harder to decode reliably.
Observation
Our first experiment used a simple Rust sentence generator. The model scored well on held-out combinations of its familiar sentence pattern, while an everyday sentence outside that pattern exposed a gap in generalization. Those results suggested it had learned the template more reliably than the underlying encoding. We changed the data and restarted from the original base model.
Mission
Teach Apertus v1.1 0.5B Base to decode Unicode Braille character by character. To isolate that skill, we restricted the task to lowercase a to z and spaces. Capitals, numbers and punctuation were excluded. Within this subset, the task is a direct mapping, not an exercise in understanding a sentence's meaning.
Approach
A small local Python generator created 17,408 random letter sequences, 4 to 48 characters long, with spaces allowed. Our existing Simple Braille library converted them into Unicode Braille. We used no external text corpus or teacher-generated prose. Every example included the full alphabet mapping, a Braille sequence, and its exact English decoding.
We trained a fresh LoRA adapter on the original pretrained base, keeping its weights frozen. Rank was 16 and alpha 32. One cycle meant one pass through the corpus: 1,088 updates, with 16 examples per batch, on one RTX 4070 Ti Super with 16 GB of VRAM. This was supervised fine-tuning only.
The run processed 4,465,930 input/output tokens, including the repeated alphabet and instructions. Of these, 259,072 were supervised target tokens. Training updates took about 14.3 minutes. The recorded run, including baseline and final evaluations and checkpoint saves, took about 19.3 minutes. Data preparation and model loading were additional.
First-cycle measurements
We held out 100 validation sequences and 100 test sequences, with no exact overlap between splits. The same 100 test inputs were evaluated before and after training, with and without the alphabet reference. No test examples were used for training.
We measured entirely correct sequences and character accuracy. Here, character accuracy is one minus character edit distance divided by reference length, floored at zero. Missing, extra and substituted characters count as errors. It is an edit-based score, not a percentage of perfectly aligned characters.
| Model | Reference | Sequences fully correct | Character accuracy |
|---|---|---|---|
| Original base | Alphabet supplied | 0% | 0.00% |
| Original base | No alphabet supplied | 0% | 0.00% |
| After one cycle | Alphabet supplied | 67% | 97.75% |
| After one cycle | No alphabet supplied | 59% | 95.39% |
The same sentence, before and after
This lowercase sentence was absent from the training, validation and test corpora. We gave the original base and the trained model exactly the same prompt within each reference condition, using greedy decoding and the same 96-token output limit. The text below is the actual model output, without corrections.
Identical Unicode Braille input
⠏⠇⠑⠁⠎⠑⠀⠕⠏⠑⠝⠀⠞⠓⠑⠀⠺⠊⠝⠙⠕⠺⠀⠃⠑⠋⠕⠗⠑⠀⠽⠕⠥⠀⠇⠑⠁⠧⠑
Expected English
please open the window before you leave
With the alphabet reference
| Original base: actual output | After one cycle: actual output |
|---|---|
| ⠏⠇⠑⠁⠎⠑⠀⠕⠏⠑⠝⠀⠞⠓⠑⠀⠺⠊⠝⠙⠕⠺⠀⠃⠑⠋⠕⠗⠑⠀⠽⠕ | please open the window before you leave |
| 0.00% character accuracy | 100.00% character accuracy |
Without the alphabet reference
| Original base: actual output | After one cycle: actual output |
|---|---|
| ⠏⠇⠑⠁⠎⠑⠀⠕⠏⠑⠝⠀⠞⠓⠑⠀⠺⠊⠝⠙⠕⠺⠀⠃⠑⠋⠕⠗⠑⠀⠽⠕ | please open teh window before you leavew before |
| 0.00% character accuracy | 74.36% character accuracy |
Progress after further training
The same test across all four checkpoints
The clearest comparison uses the exact same 100 held-out random sequences at every checkpoint. Each input contained 4 to 48 characters. All four evaluations used the native model with its adapter and no alphabet reference. Across 2,714 reference characters, the total edit distance fell from 125 to just 3.
| Checkpoint | Character accuracy | Character edits |
|---|---|---|
| First round | 95.39% | 125 |
| Round two | 96.13% | 105 |
| Round three | 97.02% | 81 |
| After overnight training | 99.89% | 3 |
This is measurable improvement on the original decoding task. The lower scores on the later everyday-sentence benchmark came from different inputs and do not demonstrate deterioration on this original test. The short test was reused to compare checkpoints, so it should not be mistaken for a fresh final benchmark.
The first cycle reached 97.75% character accuracy with the alphabet reference and 95.39% without it on unseen random sequences. Fully correct sequences rose from 0% to 67% and 59%, respectively. The comparison above preserves the actual first-cycle outputs.
After two shorter follow-up rounds, an eight-hour continuation used a new corpus of 100,000 random sequences spanning 4 to 256 characters. On the same 202 inputs evaluated before and after that continuation, character accuracy without the alphabet improved from 47.54% to 95.97%, a gain of 48.43 percentage points. This set contained 200 random sequences and two sentence probes; it was not a broad prose benchmark.
| Evaluation | Before | After |
|---|---|---|
| Mixed sequence set, 202 inputs | 47.54% | 95.97% |
| Short sentence shown above | 89.74% | 100.00% |
| Long sentence probe, 154 characters | 24.03% | 66.88% |
The short sentence became an exact translation without the reference. The longer sentence also improved, although it still needed substantial correction. These percentages show partial learning that a simple pass or fail label would hide.
Where the learning remained incomplete
A separate set of 100 newly written everyday sentences gave a more demanding view. None exactly matched the character-training or evaluation corpora. On these same new sentences, native BF16 with the adapter scored 35.14% character accuracy without the alphabet and 45.42% with it. The merged FP16 export scored 52.83% without the reference. Precision, merging and inference backend differed between the native and exported models, so that difference cannot be assigned to any one factor.
| Model | Alphabet reference | Accuracy |
|---|---|---|
| Native BF16 plus adapter | No | 35.14% |
| Native BF16 plus adapter | Yes | 45.42% |
| Merged FP16 export | No | 52.83% |
| Q4_K_M, no importance matrix | No | 0.00%* |
*The Q4 edit-based score was floored at zero because the output contained more edits than reference characters. This does not mean every output character was incorrect.
What this experiment demonstrated
A small model can acquire a useful part of the character mapping through local LoRA training. Longer training examples improved the measured ability to decode longer inputs, and the short sentence reached 100% character accuracy without assistance. The broader sentence test also showed that strong random-sequence results were not enough to establish dependable translation.
The experiment is now closed. Its outcome is a working training and evaluation process, measurable decoding progress, and a clear boundary around the result: lowercase letters and spaces, with reliable general translation still unresolved. The most useful lesson was to judge the intended use directly, rather than rely on a high score from a narrower task.