Files
el/elp/LANG-CONSOLIDATION.md
T
will.anderson a816b119e7 stage(elp): consolidate scattered lang work — full-lexicon vocabulary + profiles
Backfill ELP vocabulary from FULL lexicons (UniMorph + kaikki.org Wiktionary,
real gender/inflections) for 8 languages, 812,894 entries total, in the proven
seed-fn format matching the 18 ancient vocabularies:
  es 72,032 | fr 130,517 | de 144,692 | la 22,590 | it 193,675 | pt 115,772 |
  ro 86,504 | ca 47,112
4 of these (es fr de la) backfill ELP languages that had morphology but no
vocabulary; it/pt/ro/ca are new Romance (need morphology-*.el ports next).
Adds lang_profile_* for all 8 + reproducible generators under tests/lang-gen.
Vocab is runtime seed data (not in build manifest, like the 18 ancients);
seed-fn format validated to compile to C via elc.
2026-08-13 11:56:03 -05:00

3.9 KiB

ELP language consolidation — full-lexicon backfill (stage)

Branch: stage-elp-lang-consolidation (stage-bound; NOT the live soul :8742).

Consolidates scattered Python language-realizer work (~/Desktop/lang-realizers, ~/Desktop/lang-poetry-experiment, ~/semitic_engine) into the ELP .el structure, generating full lexicons (complete UniMorph + kaikki.org Wiktionary — real gender, real inflections) instead of the demo/curated subsets the prototypes shipped.

ELP before this branch

  • 18 classical/ancient languages fully done (vocab + morphology + tests): akk ang cop egy enm fro gez goh got grc non peo pi sa sga sux txb uga.
  • 11 modern/classical languages had morphology-<code>.el in the build manifest but no vocabulary and no lang_profile: es fr de ja ar he hi ru fi sw la.
  • The ES port (stage-elp-es-port) had a demo-scale vocabulary-es.el (~350 entries, s-expr form).

Landed on this branch (full-lexicon seed-fn format, matching the 18 ancients)

Vocabulary schema per row: [lemma, pos, form0, form1, form2, en_gloss, hint]. Files are ELP runtime seed data (loaded via the Engram at runtime), so — like all 18 classical vocabulary-*.el — they are intentionally NOT in the build manifest. Syntax validated: the chunked fn vocab_<code>_seed_pN format compiles cleanly to C via elc (correct UTF-8).

code in-ELP-morph? vocab entries verbs nouns adjs profile
es yes 72,032 6,695 48,353 16,984 yes
fr yes 130,517 7,534 77,344 45,639 yes
de yes 144,692 6,661 133,162 4,869 yes
la yes 22,590 82 13,436 9,072 yes
it no (bonus) 193,675 10,008 109,459 74,208 yes
pt no (bonus) 115,772 4,001 72,073 39,698 yes
ro no (bonus) 86,504 1,216 65,915 19,373 yes
ca no (bonus) 47,112 1,547 28,830 16,735 yes
total 812,894

Generators (reproducible): elp/tests/lang-gen/gen_elp_seed_full.py (Romance), gen_elp_seed_de_la.py (German declension + Latin case-paradigm mapping). They read the pre-built morph caches in ~/Desktop/lang-realizers/data/ (UniMorph + kaikki), which are too large to commit.

Remaining (honest)

Of the 11 ELP backfill targets, 4 are done (es fr de la). The other 7 have no full-lexicon engine yet — cannot be generated honestly without engine work:

  • ru: only a 110-entry curated Slavic subset exists; full rus.unimorph present but no morphology_ru_full productive loader. Needs a full Russian morphology module (like the Romance ones) before vocab generation.
  • ja / ko / zh: validated demo engines (~66-104 hardcoded words) in lang-poetry-experiment, Python only. Agglutinative (ja/ko) + isolating (zh) need .el engine ports + full-lexicon wiring (ja: jpn_unimorph; zh: CC-CEDICT).
  • ar / he (Semitic): template engines (16 AR / 8 HE patterns, ~6 roots) in ~/semitic_engine, Python only. Root-and-pattern; full UniMorph ara/heb present but used only for validation. Needs productive root lexicon + .el port.
  • hi (Hindi), fi (Finnish), sw (Swahili): morphology-<code>.el exists in ELP but there is NO scattered prototype and NO downloaded data for these — full-lexicon collection (UniMorph/kaikki) + generator still to do.

De/nl/sv Germanic and it/ro/ca/pt Romance verb coverage note: German verbs here are the ~6.6k caches carry; the it/ro/ca/pt bonus languages have full vocab but no morphology-<code>.el in ELP yet (Python realizer exists; .el port is the remaining engine work).

Construction coverage (separate from lexicon): French realizer was ~55%, Semitic ~3% in the prototypes — full construction coverage remains its own task.