Gaeilge AI programme · Qwen3.5-4B-Base

Irish that works as a working language.

We are building a compact open-weight Irish–English model that can write natural contemporary Irish, handle grammar and dialect, follow instructions, and eventually run locally — measured before it is claimed.

Programme statusOpen

BASE

Qwen/Qwen3.5-4B-Base · Apache 2.0

REFERENCE

Qwen3.5-4B post-trained (untouched baseline)

SIZE

~4.66B · 262K context · text-first v1

PATH

CPT → SFT → preference → GGUF

RULE

No superiority claim without shared harness

Why this matters

A language should not need permission to enter the future.

Irish has official status and living communities — yet most AI systems remain English-first and cloud-first. When tools for education, public services and everyday work are weaker in Irish, people are pushed toward English simply because the software is better. This programme exists to close that gap with evidence, not a chatbot skin.

Lower cost

Routine use without paying a cloud provider per message.

Better privacy

School, workplace and community material can stay on-device.

Offline access

Useful where connectivity is poor, expensive or unavailable.

Long-term control

Irish organisations are not locked to a foreign language priority.

Wider access

Schools and community groups can experiment without a data-centre GPU.

Local adaptation

Connect approved Irish documents for a defined institutional purpose.

Measurable ambition

Three victory levels. Zero advance claims.

The target is not a model that merely replies in Irish. It is a compact, locally deployable Irish–English model that writes natural contemporary Irish, handles mutation and idiom, retrieves and reasons over Irish information, and retains Qwen's general capability — proven under reproducible evaluation.

01

Irish-specialist state of the art

Beat fair reruns of UCCIX, Qomhrá and untouched Qwen3.5-4B under the same harness — establishing a 4B-class purpose-built open Irish model.

02

Compact open-model leadership

Beat all open models at or below 10B that can be run through the same Irish evaluation harness.

03

Public leaderboard challenge

Submit to the Irish LLM Leaderboard as an external audit — targeting the best open-weight aggregate, without treating the board as the sole development objective.

Foundation

Why Qwen3.5-4B-Base.

Qwen publishes an official pretrained-only base intended for fine-tuning — not merely an instruction model. Chat control tokens are already trained, so LoRA-style post-training does not require an expensive embedding update. The vision tower stays frozen for v1; Irish visual instruction data does not enter the first production run.

Parameters

~4.66B including vision components

Hidden size / layers

2,560 · 32 layers

Vocabulary

248,320 tied embeddings

Native context

262,144 tokens

Language coverage

201 languages and dialects reported

Licence

Apache 2.0

Method

CPT, then SFT, then preference. In that order.

We adapt the base checkpoint — not the already-aligned post-trained model — so continued pretraining does not fight prior alignment, then try to repair the damage afterward.

01

Continued pretraining

One planned bilingual Irish CPT pass on Qwen3.5-4B-Base. Parallel Irish–English material early. English and reasoning replay mandatory so general capability survives.

02

Irish-first instruction tuning

Conservative SFT from the winning CPT checkpoint — not an arbitrary later step. Every SFT checkpoint is compared to the pre-SFT model so formatting gains cannot hide language regression.

03

Fluent-speaker preference

Restrained DPO only after SFT clears regression gates. Subtle naturalness and dialect pairs require human preference; synthetic negatives cover only obvious errors.

04

Local packaging

Release adapters, merged checkpoints where appropriate, and GGUF variants tested on ordinary hardware — the claim belongs to the exact artifact users receive.

Experimental doctrine

Small first. One serious production run.

Cheap scale probes eliminate bad mixture and learning-rate decisions before the production allocation. Pilots are experiments, not candidate releases. The production run uses the winning recipe without moving the goalposts afterward.

G0

No training

Freeze baselines and validate the harness.

G1

5M tokens

Pipeline and memorisation smoke test.

G2

25M × mixtures

Choose the Irish-to-replay data mixture.

G3

50M × LRs

Choose optimizer and learning rate.

G4

100M rehearsal

Reproduce the full CPT→SFT path.

Prod

300M–800M

One planned production CPT pass.

Evaluation

GaeilgeBench before the pitch deck.

External suites — Irish-BLiMP, IRLBench, IrishQA, Silicon in Irish — stay evaluation-only. Alongside them we are building a sealed GaeilgeBench with a public development sibling. Fluent reviewers across the three major dialect regions remain the authority for naturalness.

Grammar & mutation

Irish-BLiMP + sealed minimal pairs

Natural conversation

Blind fluent-speaker win rate

Meaning preservation

Human adequacy and error taxonomy

Dialect & register

Preference by Ulster / Connacht / Munster

Cultural knowledge

Evidence-supported Irish-context QA

Reasoning in Irish

Matched Irish–English performance gap

Retrieval & abstention

Cited accuracy and false-answer rate

Safety & uncertainty

Refusal and calibration behaviour

Example promotion gates (preregistered)

Irish held-out perplexity

≥ 15% relative improvement over base

Irish-BLiMP

≥ +8 absolute points over post-trained Qwen3.5

Fluent preference

≥ 60% win rate vs post-trained Qwen3.5

Irish compliance

≥ 95% when Irish is requested

English / general regression

≤ 2 absolute points aggregate

Dialect disparity

No subgroup >10 points below aggregate

Inherited lessons

Standing on UCCIX and Qomhrá — not reinventing them.

Those projects proved Irish capability can improve through focused training. They also documented the failure modes we treat as hard requirements.

01

One CPT epoch is usually enough

Qomhrá found almost all gain in one continued-pretraining pass. We limit CPT unless held-out loss proves more is needed.

02

Replay protects English

A substantial English and general-reasoning share in the mixture protects retention better than Irish-only adaptation.

03

SFT can hurt translation

Instruction tuning can improve formatting while degrading underlying language quality. We gate every SFT checkpoint against the CPT model.

04

Fluent speakers own naturalness

LLM judges disagree with fluent speakers on subtle Irish quality. Frontier models generate candidates; speakers define linguistic truth.

05

One dialect is not enough

Reviewers and data must cover Ulster, Connacht and Munster varieties. Aggregate scores cannot hide subgroup failure.

06

Benchmarks stay sealed

IRLBench and public exams create contamination risk. Public benchmarks are evaluation-only; GaeilgeBench stays held-out.

Data & rights

Provenance first. Public availability is not permission.

Training, development and sealed evaluation live in separate stores. Candidate sources — Corpas Náisiúnta na Gaeilge, licensed Wikimedia, institutional and parallel corpora, commissioned speaker text — each receive a recorded licence and release decision before any row reaches training.

CPT mixture under test

35%

Native / edited contemporary Irish

15%

Institutional & public-service Irish

15%

Cultural / knowledge material

10%

Irish–English parallel text

5%

Conversational & code-switched Irish

15%

High-quality English / general replay

5%

Reasoning / code / math replay

Release

Artifacts people can inspect. Limitations stated plainly.

The model should answer to the Irish-language community, not only to a benchmark. Failures, dialect disagreements and negative results publish where they affect interpretation.

01LoRA adapter and merged BF16 checkpoint where appropriate

02Q4_K_M, Q5_K_M and Q8_0 GGUF variants with measured local results

03Bilingual model card stating only what was measured

04Dataset card and source-level provenance summary

05Training configuration and reproducible commands

06Public evaluation code and development set

07Known limitations, negative results and subgroup failures

08Contact path for language-quality and data concerns

Build with us

Speakers, researchers, institutions, compute partners.

LeemerLabs leads the pipeline through Born: dataset provenance, teacher orchestration, training, evaluation and local packaging. We welcome fluent reviewers across dialect regions, research partners for independent evaluation, and organisations that can help with lawful data and compute.