# ELP language consolidation — full-lexicon backfill (stage) Branch: `stage-elp-lang-consolidation` (stage-bound; NOT the live soul :8742). Consolidates scattered Python language-realizer work (`~/Desktop/lang-realizers`, `~/Desktop/lang-poetry-experiment`, `~/semitic_engine`) into the ELP `.el` structure, generating **full lexicons** (complete UniMorph + kaikki.org Wiktionary — real gender, real inflections) instead of the demo/curated subsets the prototypes shipped. ## ELP before this branch - 18 classical/ancient languages fully done (vocab + morphology + tests): akk ang cop egy enm fro gez goh got grc non peo pi sa sga sux txb uga. - 11 modern/classical languages had `morphology-.el` in the build manifest but **no vocabulary and no lang_profile**: es fr de ja ar he hi ru fi sw la. - The ES port (`stage-elp-es-port`) had a *demo-scale* vocabulary-es.el (~350 entries, s-expr form). ## Landed on this branch (full-lexicon seed-fn format, matching the 18 ancients) Vocabulary schema per row: `[lemma, pos, form0, form1, form2, en_gloss, hint]`. Files are ELP runtime **seed data** (loaded via the Engram at runtime), so — like all 18 classical `vocabulary-*.el` — they are intentionally NOT in the build manifest. Syntax validated: the chunked `fn vocab__seed_pN` format compiles cleanly to C via `elc` (correct UTF-8). | code | in-ELP-morph? | vocab entries | verbs | nouns | adjs | profile | |------|---------------|--------------:|------:|------:|-----:|---------| | es | yes | 72,032 | 6,695 | 48,353 | 16,984 | yes | | fr | yes | 130,517 | 7,534 | 77,344 | 45,639 | yes | | de | yes | 144,692 | 6,661 | 133,162 | 4,869 | yes | | la | yes | 22,590 | 82 | 13,436 | 9,072 | yes | | it | no (bonus) | 193,675 | 10,008 | 109,459 | 74,208 | yes | | pt | no (bonus) | 115,772 | 4,001 | 72,073 | 39,698 | yes | | ro | no (bonus) | 86,504 | 1,216 | 65,915 | 19,373 | yes | | ca | no (bonus) | 47,112 | 1,547 | 28,830 | 16,735 | yes | |**total**| |**812,894** | | | | | Generators (reproducible): `elp/tests/lang-gen/gen_elp_seed_full.py` (Romance), `gen_elp_seed_de_la.py` (German declension + Latin case-paradigm mapping). They read the pre-built morph caches in `~/Desktop/lang-realizers/data/` (UniMorph + kaikki), which are too large to commit. ## Remaining (honest) Of the 11 ELP backfill targets, 4 are done (es fr de la). The other 7 have **no full-lexicon engine** yet — cannot be generated honestly without engine work: - **ru**: only a 110-entry curated Slavic subset exists; full `rus.unimorph` present but no `morphology_ru_full` productive loader. Needs a full Russian morphology module (like the Romance ones) before vocab generation. - **ja / ko / zh**: validated demo engines (~66-104 hardcoded words) in `lang-poetry-experiment`, Python only. Agglutinative (ja/ko) + isolating (zh) need `.el` engine ports + full-lexicon wiring (ja: jpn_unimorph; zh: CC-CEDICT). - **ar / he (Semitic)**: template engines (16 AR / 8 HE patterns, ~6 roots) in `~/semitic_engine`, Python only. Root-and-pattern; full UniMorph ara/heb present but used only for validation. Needs productive root lexicon + `.el` port. - **hi (Hindi), fi (Finnish), sw (Swahili)**: `morphology-.el` exists in ELP but there is NO scattered prototype and NO downloaded data for these — full-lexicon collection (UniMorph/kaikki) + generator still to do. De/nl/sv Germanic and it/ro/ca/pt Romance verb coverage note: German verbs here are the ~6.6k caches carry; the it/ro/ca/pt bonus languages have full vocab but **no `morphology-.el` in ELP yet** (Python realizer exists; `.el` port is the remaining engine work). Construction coverage (separate from lexicon): French realizer was ~55%, Semitic ~3% in the prototypes — full construction coverage remains its own task.