01
Irish-specialist state of the art
Beat fair reruns of UCCIX, Qomhrá and untouched Qwen3.5-4B under the same harness — establishing a 4B-class purpose-built open Irish model.
Gaeilge AI programme · Qwen3.5-4B-Base
We are building a compact open-weight Irish–English model that can write natural contemporary Irish, handle grammar and dialect, follow instructions, and eventually run locally — measured before it is claimed.
BASE
Qwen/Qwen3.5-4B-Base · Apache 2.0
REFERENCE
Qwen3.5-4B post-trained (untouched baseline)
SIZE
~4.66B · 262K context · text-first v1
PATH
CPT → SFT → preference → GGUF
RULE
No superiority claim without shared harness
Why this matters
Irish has official status and living communities — yet most AI systems remain English-first and cloud-first. When tools for education, public services and everyday work are weaker in Irish, people are pushed toward English simply because the software is better. This programme exists to close that gap with evidence, not a chatbot skin.
Routine use without paying a cloud provider per message.
School, workplace and community material can stay on-device.
Useful where connectivity is poor, expensive or unavailable.
Irish organisations are not locked to a foreign language priority.
Schools and community groups can experiment without a data-centre GPU.
Connect approved Irish documents for a defined institutional purpose.
Measurable ambition
The target is not a model that merely replies in Irish. It is a compact, locally deployable Irish–English model that writes natural contemporary Irish, handles mutation and idiom, retrieves and reasons over Irish information, and retains Qwen's general capability — proven under reproducible evaluation.
01
Beat fair reruns of UCCIX, Qomhrá and untouched Qwen3.5-4B under the same harness — establishing a 4B-class purpose-built open Irish model.
02
Beat all open models at or below 10B that can be run through the same Irish evaluation harness.
03
Submit to the Irish LLM Leaderboard as an external audit — targeting the best open-weight aggregate, without treating the board as the sole development objective.
Foundation
Qwen publishes an official pretrained-only base intended for fine-tuning — not merely an instruction model. Chat control tokens are already trained, so LoRA-style post-training does not require an expensive embedding update. The vision tower stays frozen for v1; Irish visual instruction data does not enter the first production run.
Parameters
~4.66B including vision components
Hidden size / layers
2,560 · 32 layers
Vocabulary
248,320 tied embeddings
Native context
262,144 tokens
Language coverage
201 languages and dialects reported
Licence
Apache 2.0
Method
We adapt the base checkpoint — not the already-aligned post-trained model — so continued pretraining does not fight prior alignment, then try to repair the damage afterward.
01
One planned bilingual Irish CPT pass on Qwen3.5-4B-Base. Parallel Irish–English material early. English and reasoning replay mandatory so general capability survives.
02
Conservative SFT from the winning CPT checkpoint — not an arbitrary later step. Every SFT checkpoint is compared to the pre-SFT model so formatting gains cannot hide language regression.
03
Restrained DPO only after SFT clears regression gates. Subtle naturalness and dialect pairs require human preference; synthetic negatives cover only obvious errors.
04
Release adapters, merged checkpoints where appropriate, and GGUF variants tested on ordinary hardware — the claim belongs to the exact artifact users receive.
Experimental doctrine
Cheap scale probes eliminate bad mixture and learning-rate decisions before the production allocation. Pilots are experiments, not candidate releases. The production run uses the winning recipe without moving the goalposts afterward.
G0
No training
Freeze baselines and validate the harness.
G1
5M tokens
Pipeline and memorisation smoke test.
G2
25M × mixtures
Choose the Irish-to-replay data mixture.
G3
50M × LRs
Choose optimizer and learning rate.
G4
100M rehearsal
Reproduce the full CPT→SFT path.
Prod
300M–800M
One planned production CPT pass.
Evaluation
External suites — Irish-BLiMP, IRLBench, IrishQA, Silicon in Irish — stay evaluation-only. Alongside them we are building a sealed GaeilgeBench with a public development sibling. Fluent reviewers across the three major dialect regions remain the authority for naturalness.
Grammar & mutation
Irish-BLiMP + sealed minimal pairs
Natural conversation
Blind fluent-speaker win rate
Meaning preservation
Human adequacy and error taxonomy
Dialect & register
Preference by Ulster / Connacht / Munster
Cultural knowledge
Evidence-supported Irish-context QA
Reasoning in Irish
Matched Irish–English performance gap
Retrieval & abstention
Cited accuracy and false-answer rate
Safety & uncertainty
Refusal and calibration behaviour
Example promotion gates (preregistered)
Irish held-out perplexity
≥ 15% relative improvement over base
Irish-BLiMP
≥ +8 absolute points over post-trained Qwen3.5
Fluent preference
≥ 60% win rate vs post-trained Qwen3.5
Irish compliance
≥ 95% when Irish is requested
English / general regression
≤ 2 absolute points aggregate
Dialect disparity
No subgroup >10 points below aggregate
Inherited lessons
Those projects proved Irish capability can improve through focused training. They also documented the failure modes we treat as hard requirements.
01
Qomhrá found almost all gain in one continued-pretraining pass. We limit CPT unless held-out loss proves more is needed.
02
A substantial English and general-reasoning share in the mixture protects retention better than Irish-only adaptation.
03
Instruction tuning can improve formatting while degrading underlying language quality. We gate every SFT checkpoint against the CPT model.
04
LLM judges disagree with fluent speakers on subtle Irish quality. Frontier models generate candidates; speakers define linguistic truth.
05
Reviewers and data must cover Ulster, Connacht and Munster varieties. Aggregate scores cannot hide subgroup failure.
06
IRLBench and public exams create contamination risk. Public benchmarks are evaluation-only; GaeilgeBench stays held-out.
Data & rights
Training, development and sealed evaluation live in separate stores. Candidate sources — Corpas Náisiúnta na Gaeilge, licensed Wikimedia, institutional and parallel corpora, commissioned speaker text — each receive a recorded licence and release decision before any row reaches training.
CPT mixture under test
35%
Native / edited contemporary Irish
15%
Institutional & public-service Irish
15%
Cultural / knowledge material
10%
Irish–English parallel text
5%
Conversational & code-switched Irish
15%
High-quality English / general replay
5%
Reasoning / code / math replay
Release
The model should answer to the Irish-language community, not only to a benchmark. Failures, dialect disagreements and negative results publish where they affect interpretation.
01LoRA adapter and merged BF16 checkpoint where appropriate
02Q4_K_M, Q5_K_M and Q8_0 GGUF variants with measured local results
03Bilingual model card stating only what was measured
04Dataset card and source-level provenance summary
05Training configuration and reproducible commands
06Public evaluation code and development set
07Known limitations, negative results and subgroup failures
08Contact path for language-quality and data concerns
Build with us
LeemerLabs leads the pipeline through Born: dataset provenance, teacher orchestration, training, evaluation and local packaging. We welcome fluent reviewers across dialect regions, research partners for independent evaluation, and organisations that can help with lawful data and compute.