a816b119e7
Backfill ELP vocabulary from FULL lexicons (UniMorph + kaikki.org Wiktionary, real gender/inflections) for 8 languages, 812,894 entries total, in the proven seed-fn format matching the 18 ancient vocabularies: es 72,032 | fr 130,517 | de 144,692 | la 22,590 | it 193,675 | pt 115,772 | ro 86,504 | ca 47,112 4 of these (es fr de la) backfill ELP languages that had morphology but no vocabulary; it/pt/ro/ca are new Romance (need morphology-*.el ports next). Adds lang_profile_* for all 8 + reproducible generators under tests/lang-gen. Vocab is runtime seed data (not in build manifest, like the 18 ancients); seed-fn format validated to compile to C via elc.
66 lines
3.9 KiB
Markdown
66 lines
3.9 KiB
Markdown
# ELP language consolidation — full-lexicon backfill (stage)
|
|
|
|
Branch: `stage-elp-lang-consolidation` (stage-bound; NOT the live soul :8742).
|
|
|
|
Consolidates scattered Python language-realizer work (`~/Desktop/lang-realizers`,
|
|
`~/Desktop/lang-poetry-experiment`, `~/semitic_engine`) into the ELP `.el`
|
|
structure, generating **full lexicons** (complete UniMorph + kaikki.org
|
|
Wiktionary — real gender, real inflections) instead of the demo/curated subsets
|
|
the prototypes shipped.
|
|
|
|
## ELP before this branch
|
|
- 18 classical/ancient languages fully done (vocab + morphology + tests):
|
|
akk ang cop egy enm fro gez goh got grc non peo pi sa sga sux txb uga.
|
|
- 11 modern/classical languages had `morphology-<code>.el` in the build manifest
|
|
but **no vocabulary and no lang_profile**: es fr de ja ar he hi ru fi sw la.
|
|
- The ES port (`stage-elp-es-port`) had a *demo-scale* vocabulary-es.el (~350
|
|
entries, s-expr form).
|
|
|
|
## Landed on this branch (full-lexicon seed-fn format, matching the 18 ancients)
|
|
Vocabulary schema per row: `[lemma, pos, form0, form1, form2, en_gloss, hint]`.
|
|
Files are ELP runtime **seed data** (loaded via the Engram at runtime), so — like
|
|
all 18 classical `vocabulary-*.el` — they are intentionally NOT in the build
|
|
manifest. Syntax validated: the chunked `fn vocab_<code>_seed_pN` format
|
|
compiles cleanly to C via `elc` (correct UTF-8).
|
|
|
|
| code | in-ELP-morph? | vocab entries | verbs | nouns | adjs | profile |
|
|
|------|---------------|--------------:|------:|------:|-----:|---------|
|
|
| es | yes | 72,032 | 6,695 | 48,353 | 16,984 | yes |
|
|
| fr | yes | 130,517 | 7,534 | 77,344 | 45,639 | yes |
|
|
| de | yes | 144,692 | 6,661 | 133,162 | 4,869 | yes |
|
|
| la | yes | 22,590 | 82 | 13,436 | 9,072 | yes |
|
|
| it | no (bonus) | 193,675 | 10,008 | 109,459 | 74,208 | yes |
|
|
| pt | no (bonus) | 115,772 | 4,001 | 72,073 | 39,698 | yes |
|
|
| ro | no (bonus) | 86,504 | 1,216 | 65,915 | 19,373 | yes |
|
|
| ca | no (bonus) | 47,112 | 1,547 | 28,830 | 16,735 | yes |
|
|
|**total**| |**812,894** | | | | |
|
|
|
|
Generators (reproducible): `elp/tests/lang-gen/gen_elp_seed_full.py` (Romance),
|
|
`gen_elp_seed_de_la.py` (German declension + Latin case-paradigm mapping). They
|
|
read the pre-built morph caches in `~/Desktop/lang-realizers/data/` (UniMorph +
|
|
kaikki), which are too large to commit.
|
|
|
|
## Remaining (honest)
|
|
Of the 11 ELP backfill targets, 4 are done (es fr de la). The other 7 have **no
|
|
full-lexicon engine** yet — cannot be generated honestly without engine work:
|
|
- **ru**: only a 110-entry curated Slavic subset exists; full `rus.unimorph`
|
|
present but no `morphology_ru_full` productive loader. Needs a full Russian
|
|
morphology module (like the Romance ones) before vocab generation.
|
|
- **ja / ko / zh**: validated demo engines (~66-104 hardcoded words) in
|
|
`lang-poetry-experiment`, Python only. Agglutinative (ja/ko) + isolating (zh)
|
|
need `.el` engine ports + full-lexicon wiring (ja: jpn_unimorph; zh: CC-CEDICT).
|
|
- **ar / he (Semitic)**: template engines (16 AR / 8 HE patterns, ~6 roots) in
|
|
`~/semitic_engine`, Python only. Root-and-pattern; full UniMorph ara/heb
|
|
present but used only for validation. Needs productive root lexicon + `.el` port.
|
|
- **hi (Hindi), fi (Finnish), sw (Swahili)**: `morphology-<code>.el` exists in
|
|
ELP but there is NO scattered prototype and NO downloaded data for these —
|
|
full-lexicon collection (UniMorph/kaikki) + generator still to do.
|
|
|
|
De/nl/sv Germanic and it/ro/ca/pt Romance verb coverage note: German verbs here
|
|
are the ~6.6k caches carry; the it/ro/ca/pt bonus languages have full vocab but
|
|
**no `morphology-<code>.el` in ELP yet** (Python realizer exists; `.el` port is
|
|
the remaining engine work).
|
|
|
|
Construction coverage (separate from lexicon): French realizer was ~55%,
|
|
Semitic ~3% in the prototypes — full construction coverage remains its own task.
|