Backfill ELP vocabulary from FULL lexicons (UniMorph + kaikki.org Wiktionary, real gender/inflections) for 8 languages, 812,894 entries total, in the proven seed-fn format matching the 18 ancient vocabularies: es 72,032 | fr 130,517 | de 144,692 | la 22,590 | it 193,675 | pt 115,772 | ro 86,504 | ca 47,112 4 of these (es fr de la) backfill ELP languages that had morphology but no vocabulary; it/pt/ro/ca are new Romance (need morphology-*.el ports next). Adds lang_profile_* for all 8 + reproducible generators under tests/lang-gen. Vocab is runtime seed data (not in build manifest, like the 18 ancients); seed-fn format validated to compile to C via elc.
3.9 KiB
ELP language consolidation — full-lexicon backfill (stage)
Branch: stage-elp-lang-consolidation (stage-bound; NOT the live soul :8742).
Consolidates scattered Python language-realizer work (~/Desktop/lang-realizers,
~/Desktop/lang-poetry-experiment, ~/semitic_engine) into the ELP .el
structure, generating full lexicons (complete UniMorph + kaikki.org
Wiktionary — real gender, real inflections) instead of the demo/curated subsets
the prototypes shipped.
ELP before this branch
- 18 classical/ancient languages fully done (vocab + morphology + tests): akk ang cop egy enm fro gez goh got grc non peo pi sa sga sux txb uga.
- 11 modern/classical languages had
morphology-<code>.elin the build manifest but no vocabulary and no lang_profile: es fr de ja ar he hi ru fi sw la. - The ES port (
stage-elp-es-port) had a demo-scale vocabulary-es.el (~350 entries, s-expr form).
Landed on this branch (full-lexicon seed-fn format, matching the 18 ancients)
Vocabulary schema per row: [lemma, pos, form0, form1, form2, en_gloss, hint].
Files are ELP runtime seed data (loaded via the Engram at runtime), so — like
all 18 classical vocabulary-*.el — they are intentionally NOT in the build
manifest. Syntax validated: the chunked fn vocab_<code>_seed_pN format
compiles cleanly to C via elc (correct UTF-8).
| code | in-ELP-morph? | vocab entries | verbs | nouns | adjs | profile |
|---|---|---|---|---|---|---|
| es | yes | 72,032 | 6,695 | 48,353 | 16,984 | yes |
| fr | yes | 130,517 | 7,534 | 77,344 | 45,639 | yes |
| de | yes | 144,692 | 6,661 | 133,162 | 4,869 | yes |
| la | yes | 22,590 | 82 | 13,436 | 9,072 | yes |
| it | no (bonus) | 193,675 | 10,008 | 109,459 | 74,208 | yes |
| pt | no (bonus) | 115,772 | 4,001 | 72,073 | 39,698 | yes |
| ro | no (bonus) | 86,504 | 1,216 | 65,915 | 19,373 | yes |
| ca | no (bonus) | 47,112 | 1,547 | 28,830 | 16,735 | yes |
| total | 812,894 |
Generators (reproducible): elp/tests/lang-gen/gen_elp_seed_full.py (Romance),
gen_elp_seed_de_la.py (German declension + Latin case-paradigm mapping). They
read the pre-built morph caches in ~/Desktop/lang-realizers/data/ (UniMorph +
kaikki), which are too large to commit.
Remaining (honest)
Of the 11 ELP backfill targets, 4 are done (es fr de la). The other 7 have no full-lexicon engine yet — cannot be generated honestly without engine work:
- ru: only a 110-entry curated Slavic subset exists; full
rus.unimorphpresent but nomorphology_ru_fullproductive loader. Needs a full Russian morphology module (like the Romance ones) before vocab generation. - ja / ko / zh: validated demo engines (~66-104 hardcoded words) in
lang-poetry-experiment, Python only. Agglutinative (ja/ko) + isolating (zh) need.elengine ports + full-lexicon wiring (ja: jpn_unimorph; zh: CC-CEDICT). - ar / he (Semitic): template engines (16 AR / 8 HE patterns, ~6 roots) in
~/semitic_engine, Python only. Root-and-pattern; full UniMorph ara/heb present but used only for validation. Needs productive root lexicon +.elport. - hi (Hindi), fi (Finnish), sw (Swahili):
morphology-<code>.elexists in ELP but there is NO scattered prototype and NO downloaded data for these — full-lexicon collection (UniMorph/kaikki) + generator still to do.
De/nl/sv Germanic and it/ro/ca/pt Romance verb coverage note: German verbs here
are the ~6.6k caches carry; the it/ro/ca/pt bonus languages have full vocab but
no morphology-<code>.el in ELP yet (Python realizer exists; .el port is
the remaining engine work).
Construction coverage (separate from lexicon): French realizer was ~55%, Semitic ~3% in the prototypes — full construction coverage remains its own task.