Back to The Quiet Wire

A local AI experiment

Training a Small Model to Decode Unicode Braille

Using LoRA on a single consumer GPU, character accuracy rose from 95.39% to 99.89% on the same 100 short test sequences, without an alphabet reference. Longer everyday sentences remained harder to decode reliably.

AI in Unicode Braille: ⠠⠠⠁⠊
“AI” in Braille. This experiment trains lowercase letters and spaces only.

Observation

Our first experiment used a simple Rust sentence generator. The model scored well on held-out combinations of its familiar sentence pattern, while an everyday sentence outside that pattern exposed a gap in generalization. Those results suggested it had learned the template more reliably than the underlying encoding. We changed the data and restarted from the original base model.

Mission

Teach Apertus v1.1 0.5B Base to decode Unicode Braille character by character. To isolate that skill, we restricted the task to lowercase a to z and spaces. Capitals, numbers and punctuation were excluded. Within this subset, the task is a direct mapping, not an exercise in understanding a sentence's meaning.

Approach

A small local Python generator created 17,408 random letter sequences, 4 to 48 characters long, with spaces allowed. Our existing Simple Braille library converted them into Unicode Braille. We used no external text corpus or teacher-generated prose. Every example included the full alphabet mapping, a Braille sequence, and its exact English decoding.

We trained a fresh LoRA adapter on the original pretrained base, keeping its weights frozen. Rank was 16 and alpha 32. One cycle meant one pass through the corpus: 1,088 updates, with 16 examples per batch, on one RTX 4070 Ti Super with 16 GB of VRAM. This was supervised fine-tuning only.

The run processed 4,465,930 input/output tokens, including the repeated alphabet and instructions. Of these, 259,072 were supervised target tokens. Training updates took about 14.3 minutes. The recorded run, including baseline and final evaluations and checkpoint saves, took about 19.3 minutes. Data preparation and model loading were additional.

First-cycle measurements

We held out 100 validation sequences and 100 test sequences, with no exact overlap between splits. The same 100 test inputs were evaluated before and after training, with and without the alphabet reference. No test examples were used for training.

We measured entirely correct sequences and character accuracy. Here, character accuracy is one minus character edit distance divided by reference length, floored at zero. Missing, extra and substituted characters count as errors. It is an edit-based score, not a percentage of perfectly aligned characters.

100 unseen random sequences, 2,714 reference characters
ModelReferenceSequences fully correctCharacter accuracy
Original baseAlphabet supplied0%0.00%
Original baseNo alphabet supplied0%0.00%
After one cycleAlphabet supplied67%97.75%
After one cycleNo alphabet supplied59%95.39%

The same sentence, before and after

This lowercase sentence was absent from the training, validation and test corpora. We gave the original base and the trained model exactly the same prompt within each reference condition, using greedy decoding and the same 96-token output limit. The text below is the actual model output, without corrections.

Identical Unicode Braille input

⠏⠇⠑⠁⠎⠑⠀⠕⠏⠑⠝⠀⠞⠓⠑⠀⠺⠊⠝⠙⠕⠺⠀⠃⠑⠋⠕⠗⠑⠀⠽⠕⠥⠀⠇⠑⠁⠧⠑

Expected English

please open the window before you leave

With the alphabet reference

Original base: actual outputAfter one cycle: actual output
⠏⠇⠑⠁⠎⠑⠀⠕⠏⠑⠝⠀⠞⠓⠑⠀⠺⠊⠝⠙⠕⠺⠀⠃⠑⠋⠕⠗⠑⠀⠽⠕please open the window before you leave
0.00% character accuracy100.00% character accuracy

Without the alphabet reference

Original base: actual outputAfter one cycle: actual output
⠏⠇⠑⠁⠎⠑⠀⠕⠏⠑⠝⠀⠞⠓⠑⠀⠺⠊⠝⠙⠕⠺⠀⠃⠑⠋⠕⠗⠑⠀⠽⠕please open teh window before you leavew before
0.00% character accuracy74.36% character accuracy

Progress after further training

The same test across all four checkpoints

The clearest comparison uses the exact same 100 held-out random sequences at every checkpoint. Each input contained 4 to 48 characters. All four evaluations used the native model with its adapter and no alphabet reference. Across 2,714 reference characters, the total edit distance fell from 125 to just 3.

Identical 100 short inputs, no alphabet reference
CheckpointCharacter accuracyCharacter edits
First round95.39%125
Round two96.13%105
Round three97.02%81
After overnight training99.89%3

This is measurable improvement on the original decoding task. The lower scores on the later everyday-sentence benchmark came from different inputs and do not demonstrate deterioration on this original test. The short test was reused to compare checkpoints, so it should not be mistaken for a fresh final benchmark.

The first cycle reached 97.75% character accuracy with the alphabet reference and 95.39% without it on unseen random sequences. Fully correct sequences rose from 0% to 67% and 59%, respectively. The comparison above preserves the actual first-cycle outputs.

After two shorter follow-up rounds, an eight-hour continuation used a new corpus of 100,000 random sequences spanning 4 to 256 characters. On the same 202 inputs evaluated before and after that continuation, character accuracy without the alphabet improved from 47.54% to 95.97%, a gain of 48.43 percentage points. This set contained 200 random sequences and two sentence probes; it was not a broad prose benchmark.

Eight-hour continuation: no alphabet reference
EvaluationBeforeAfter
Mixed sequence set, 202 inputs47.54%95.97%
Short sentence shown above89.74%100.00%
Long sentence probe, 154 characters24.03%66.88%

The short sentence became an exact translation without the reference. The longer sentence also improved, although it still needed substantial correction. These percentages show partial learning that a simple pass or fail label would hide.

Where the learning remained incomplete

A separate set of 100 newly written everyday sentences gave a more demanding view. None exactly matched the character-training or evaluation corpora. On these same new sentences, native BF16 with the adapter scored 35.14% character accuracy without the alphabet and 45.42% with it. The merged FP16 export scored 52.83% without the reference. Precision, merging and inference backend differed between the native and exported models, so that difference cannot be assigned to any one factor.

100 new everyday sentences, character accuracy
ModelAlphabet referenceAccuracy
Native BF16 plus adapterNo35.14%
Native BF16 plus adapterYes45.42%
Merged FP16 exportNo52.83%
Q4_K_M, no importance matrixNo0.00%*

*The Q4 edit-based score was floored at zero because the output contained more edits than reference characters. This does not mean every output character was incorrect.

What this experiment demonstrated

A small model can acquire a useful part of the character mapping through local LoRA training. Longer training examples improved the measured ability to decode longer inputs, and the short sentence reached 100% character accuracy without assistance. The broader sentence test also showed that strong random-sequence results were not enough to establish dependable translation.

The experiment is now closed. Its outcome is a working training and evaluation process, measurable decoding progress, and a clear boundary around the result: lowercase letters and spaces, with reliable general translation still unresolved. The most useful lesson was to judge the intended use directly, rather than rely on a high score from a narrower task.