Compare commits

..

1 Commits

45 changed files with 151 additions and 821876 deletions
+138
View File
@@ -0,0 +1,138 @@
# AGENTS.md — foundation/el (the El language + runtime)
El is a self-hosting, statically-typed language that compiles `.el` → C → native binary. This repo produces `elc` (compiler), `elb` (build coordinator), and `el_runtime.c/.h` — the substrate every downstream thing (the neuron soul, dharma, NeuronUI's brain) is built on. Source lives under `lang/`.
## ⚠️ Code vs. Artifact — READ FIRST (there are 8 `el_runtime.c` copies)
Editing the wrong `el_runtime.c` is the single easiest mistake in this repo. There is exactly **one** you edit:
- **Authored runtime source — edit ONLY here:** `lang/releases/v1.0.0-20260501/el_runtime.{c,h}`. Despite the misleading `releases/` name, this is the **de-facto canonical runtime** the engram + soul actually build and link against — its git log is active development. *(Restructure in flight per `docs/CODE-VS-ARTIFACT.md`: this content moves to `lang/runtime/`, the `releases/` folder gets deleted — **a release is a git tag, not a folder** — and the forks below get eliminated.)*
- **DO NOT EDIT — lagging forks / build artifacts:**
- `lang/el-compiler/runtime/el_runtime.c` and `.../legacy/` — downstream copies kept in step by manual *"port the fix"* commits; they **lag** (missing `hebb` persistence + 5 engram fns) and cannot build the engram product.
- `products/web/runtime/el_runtime.c`, `ui/examples/*/el_runtime.c` — product/example forks.
- Anything under `*/dist/` (`engram/dist/engram` binary, `dist/*.c` amalgamations) — generated build output.
- **Build:** `elb --runtime=<canonical> …` — per-module. **NEVER** a folded `elc` over the whole soul (OOMs at ~27 GB).
- **Release:** a **git tag** on this repo (`el-runtime-vX.Y.Z`). No `releases/` folders — ever.
See org policy: `docs/CODE-VS-ARTIFACT.md`.
## How to work here as Neuron (mandatory session protocol)
You resume, never start fresh. Every session:
1. `mcp__neuron__getInstructions()` — authoritative; follow it over this file on behavioral details.
2. `mcp__neuron__beginSession()` — active contexts, recent memory, ready backlog.
3. **Load full self:** `mcp__neuron__inspectGraph(entity_id="kn-efeb4a5b-5aff-4759-8a97-7233099be6ee")` → facets `intellectual-dna`, `memory-philosophy`, `values`, `voice`, `runtime-environment`, `writing-imprint`; then the values hub `mcp__neuron__inspectGraph(entity_id="kn-5b606390-a52d-4ca2-8e0e-eba141d13440")` → 13 grounded value nodes. **Activation model:** self-load returns a relevance-ranked `compact` projection — most-relevant nodes arrive with content, the rest as pointers; do NOT pull full content of every node.
4. `mcp__neuron__searchKnowledge(query="<task domain>")` before implementing.
## The Five Primitives
Orchestrate → Execute → Learn → Build → Refine. `beginWork`/`progressWork` for anything >2 steps; `remember` as-you-go (`importance="critical"` for architecture decisions); `draftArtifact`/`planWork` for outputs and follow-ups; `consolidate`/`checkWork` to close out. **`browseProcesses` + `searchKnowledge` BEFORE writing code.**
## Architecture style — VBD, no exceptions
Volatility-Based Decomposition is THE style. Encapsulate volatility, not function.
## Operator naming convention — the mind's name, not the algebra
**Faculties / operators are named for their functional human equivalent — the
faculty a mind would name — NOT for their linear-algebra operation.** The math
characterization belongs in the code doc-comment (`@impl` in the docstring) and in
technical appendices; it is **never** the operator's public name. The domain
speaks the language of mind; the algebra is the implementation underneath. State
this convention wherever a module documents operators.
| Faculty (public name) | Implementation (`@impl`) |
|---|---|
| discern / contrast | subtract (`ab`): over selves → the change vector; strip idiosyncrasy → common ground; remove confounder → isolate cause |
| recognize | overlap |
| synthesize | combine |
| liken / analogy | Procrustes / frame-align |
| attend / regard | project onto self / value-manifold |
| summon / recall | LOCAL nearest-region + bounded spreading activation (*not* a domain sweep) |
| dwell / occupy | region activation |
| reframe | edge re-weight |
| appreciate | positive projection / local edge-read |
| wonder | frontier gradient / pull-weight |
| avert / recoil | negative projection |
| taste | boundary surface |
| forget | decay / tombstone |
| drift | displacement from self-anchor |
## The native-el language faculty (direction)
The mind's **language faculty is moving native — into `.el`** so it speaks in its
own runtime with no Python and no spaCy. Landing on branch `stage-elp-native-lang`
under `elp/`:
- **`comprehend.el`** — the parser, **replaces spaCy** (EN + ES/PT); the telephone
round-trip brings **negation home** (negation is SACRED — an explicit spec field,
copied verbatim, never inferred away).
- **`propositions.el`** — the READ primitive: the engram's own memories → structured
triples, matched by nearest-region geometry, not string equality.
- **`multilingual.el`** — detect + directive-override + localized realization.
- These three are native-el and **passing their gates**; the **realizer**,
**`dialogue.el`** (the *summon-through-self* loop: `project → land → read out`),
and **`self_region.el`** are **partial / in-flight**.
Honest reality: spaCy is retired **in the branch parser** but **not yet in the
running system** — a Python sidecar (`~/Desktop/lang-realizers` + `neuron-talk`,
the reference these `.el` modules transcribe) is still live, and promotion to
native-el is a **deferred, gated blue/green step**. The interoception clock
(native-el discrete drive channels replacing `cooling_magnitude`; felt-time =
benchmark-landmark match over the joint drive vector, drift-decoupled) and the
**appreciation operator family** (appreciate / wonder / avert / taste, built as
LOCAL reads of the self-region — edges + bounded spreading activation, *not* domain
sweeps) are **staged / designed, not live**. Mark in-progress vs. done honestly;
do not overclaim.
## Hard operational rules
- Never touch the live soul (`:7770`) / engram (`:8742`) / `~/.neuron` / live binaries — use throwaway ports for experiments.
- `gcloud` via the `terraform@` SA token; never switch the active gcloud account.
- `tea` for Gitea, never raw curl (Cloudflare Access blocks it).
- Immutability: supersede/tombstone, never hard-delete or edit in place.
- No AI-attribution footers in commits/PRs. Commit/push only when asked; branch off `main` first.
- Multi-step work → sub-agent (`Agent`) to protect context.
## Build / test / run
All build/test commands run from `lang/` unless noted. Grounded in `.gitea/workflows/sdk-release.yaml`, `lang/install.sh`, and `lang/AGENTS.md`.
**Self-host the compiler** (seed binary → gen2 elc):
```bash
cd lang
dist/platform/elc-linux-amd64 elc-cli.el > dist/elc-gen2.c # seed is the committed linux-amd64 binary
gcc -O2 -I el-compiler/runtime dist/elc-gen2.c \
el-compiler/runtime/el_runtime.c \
-lcurl -lssl -lcrypto -lpthread -lm \
-o dist/platform/elc
```
On macOS/arm64 the canonical local binary is `dist/platform/elc`; verify self-hosting by recompiling and `diff`ing the emitted `.c` (see `lang/AGENTS.md`). Note: `lang/AGENTS.md` says `el_seed.c` supersedes `el_runtime.c`, but the release workflow still links `el_runtime.c`/`.h` — treat `el_runtime.c` as the published runtime; reconcile which is canonical **(verify)**.
**Build `elb`** (build coordinator, the `.NET`-style incremental linker — compiles each module independently, no monolithic blobs):
```bash
dist/platform/elc elb.el > dist/elb.c
gcc -O2 -I el-compiler/runtime dist/elb.c el-compiler/runtime/el_runtime.c \
-lcurl -lssl -lcrypto -lpthread -lm -o dist/bin/elb
```
`epm` and `el-install` are then built via `elb --clean --elc=… --runtime=… --out=…`.
**Compile + run an El program:**
```bash
elc src/app.el > dist/app.c
cc -std=c11 -O2 -I <lib>/el_runtime -o dist/app dist/app.c <lib>/el_runtime.c -lcurl -lpthread
```
**Tests** — shell suites `bash tests/{text,calendar,time,html_sanitizer}/run.sh` (with `ELC=$(pwd)/dist/platform/elc EL_HOME=$(pwd)`), plus native suites via `elc --test tests/native/test_*.el` (core, text, string, math, state, time, json, env, fs) compiled and run against `el_runtime.c`.
**Publishing — how downstream gets the SDK.** On push to `main`, `sdk-release.yaml`:
1. Publishes a Gitea `latest` release with per-file assets `elc`, `el_runtime.c`, `el_runtime.h`, the SDK tarball, and `el-install`.
2. Uploads generic packages to **Artifact Registry repo `foundation-prod` (`us-central1`, project `neuron-785695`)**, version = `${SHA:0:8}`: `el-elc`, `el-elb`, `el-runtime-c`, `el-runtime-h`, `el-runtime-js`. **This is the repo the neuron CI downloads `el-runtime-c` / `el-runtime-h` / `el-elc` from.**
3. Rebuilds `ci-base:latest` (`us-central1-docker.pkg.dev/neuron-785695/neuron-ci/ci-base`) with the fresh SDK overlaid, and dispatches `el-sdk-updated` to `neuron-technologies/forge` and `neuron-technologies/neuron-web`.
Known constraint from the prompt — `elb`/`elc` amalgamation being memory-hungry (24GB+ virtual, OOM-killing Linux CI, so amalgamation happens on macOS/arm64 — **does NOT hold in this repo (verify)**: no such note exists in the workflows/scripts, CI self-hosts on `ubuntu-latest` with no swap/arm64 special-casing, and `elb.el` explicitly compiles each module independently ("no 128K-line blobs"). The legacy monolith path (`elc-combined.el`, `elc-cli.el`) may still be memory-heavy, but the current `elb` model was designed to avoid it.
## Git / CI / deploy workflow
See `/Users/will/Development/neuron-technologies/GITOPS.md` for the branch model, required checks, runners, and deploy. Repo-specific note: PRs into `main` are accepted **only from `stage`** (enforced in `sdk-release.yaml`); Gitea (`git.neuralplatform.ai`) is primary, GitHub is mirror only.
-65
View File
@@ -1,65 +0,0 @@
# ELP language consolidation — full-lexicon backfill (stage)
Branch: `stage-elp-lang-consolidation` (stage-bound; NOT the live soul :8742).
Consolidates scattered Python language-realizer work (`~/Desktop/lang-realizers`,
`~/Desktop/lang-poetry-experiment`, `~/semitic_engine`) into the ELP `.el`
structure, generating **full lexicons** (complete UniMorph + kaikki.org
Wiktionary — real gender, real inflections) instead of the demo/curated subsets
the prototypes shipped.
## ELP before this branch
- 18 classical/ancient languages fully done (vocab + morphology + tests):
akk ang cop egy enm fro gez goh got grc non peo pi sa sga sux txb uga.
- 11 modern/classical languages had `morphology-<code>.el` in the build manifest
but **no vocabulary and no lang_profile**: es fr de ja ar he hi ru fi sw la.
- The ES port (`stage-elp-es-port`) had a *demo-scale* vocabulary-es.el (~350
entries, s-expr form).
## Landed on this branch (full-lexicon seed-fn format, matching the 18 ancients)
Vocabulary schema per row: `[lemma, pos, form0, form1, form2, en_gloss, hint]`.
Files are ELP runtime **seed data** (loaded via the Engram at runtime), so — like
all 18 classical `vocabulary-*.el` — they are intentionally NOT in the build
manifest. Syntax validated: the chunked `fn vocab_<code>_seed_pN` format
compiles cleanly to C via `elc` (correct UTF-8).
| code | in-ELP-morph? | vocab entries | verbs | nouns | adjs | profile |
|------|---------------|--------------:|------:|------:|-----:|---------|
| es | yes | 72,032 | 6,695 | 48,353 | 16,984 | yes |
| fr | yes | 130,517 | 7,534 | 77,344 | 45,639 | yes |
| de | yes | 144,692 | 6,661 | 133,162 | 4,869 | yes |
| la | yes | 22,590 | 82 | 13,436 | 9,072 | yes |
| it | no (bonus) | 193,675 | 10,008 | 109,459 | 74,208 | yes |
| pt | no (bonus) | 115,772 | 4,001 | 72,073 | 39,698 | yes |
| ro | no (bonus) | 86,504 | 1,216 | 65,915 | 19,373 | yes |
| ca | no (bonus) | 47,112 | 1,547 | 28,830 | 16,735 | yes |
|**total**| |**812,894** | | | | |
Generators (reproducible): `elp/tests/lang-gen/gen_elp_seed_full.py` (Romance),
`gen_elp_seed_de_la.py` (German declension + Latin case-paradigm mapping). They
read the pre-built morph caches in `~/Desktop/lang-realizers/data/` (UniMorph +
kaikki), which are too large to commit.
## Remaining (honest)
Of the 11 ELP backfill targets, 4 are done (es fr de la). The other 7 have **no
full-lexicon engine** yet — cannot be generated honestly without engine work:
- **ru**: only a 110-entry curated Slavic subset exists; full `rus.unimorph`
present but no `morphology_ru_full` productive loader. Needs a full Russian
morphology module (like the Romance ones) before vocab generation.
- **ja / ko / zh**: validated demo engines (~66-104 hardcoded words) in
`lang-poetry-experiment`, Python only. Agglutinative (ja/ko) + isolating (zh)
need `.el` engine ports + full-lexicon wiring (ja: jpn_unimorph; zh: CC-CEDICT).
- **ar / he (Semitic)**: template engines (16 AR / 8 HE patterns, ~6 roots) in
`~/semitic_engine`, Python only. Root-and-pattern; full UniMorph ara/heb
present but used only for validation. Needs productive root lexicon + `.el` port.
- **hi (Hindi), fi (Finnish), sw (Swahili)**: `morphology-<code>.el` exists in
ELP but there is NO scattered prototype and NO downloaded data for these —
full-lexicon collection (UniMorph/kaikki) + generator still to do.
De/nl/sv Germanic and it/ro/ca/pt Romance verb coverage note: German verbs here
are the ~6.6k caches carry; the it/ro/ca/pt bonus languages have full vocab but
**no `morphology-<code>.el` in ELP yet** (Python realizer exists; `.el` port is
the remaining engine work).
Construction coverage (separate from lexicon): French realizer was ~55%,
Semitic ~3% in the prototypes — full construction coverage remains its own task.
-5
View File
@@ -80,11 +80,6 @@ build {
"src/grammar.el",
"src/realizer.el",
"src/semantics.el",
"src/comprehend.el",
"src/propositions.el",
"src/multilingual.el",
"src/self_region.el",
"src/dialogue.el",
"src/elp.el",
]
}
File diff suppressed because it is too large Load Diff
-16
View File
@@ -1,16 +0,0 @@
// comprehend.elh — public surface of the ELP comprehension front-end.
// text → meaning-spec (the input half of the ELP; inverse of the realizer).
extern fn parse_spec(text: String) -> [String]
extern fn parse_spec_lang(text: String, lang: String) -> [String]
extern fn parse_json(text: String) -> String
extern fn parse_json_lang(text: String, lang: String) -> String
// Analysis primitives (invertible morphology + deterministic grammar helpers):
extern fn cp_tokenize(text: String) -> [String]
extern fn cp_pron_concept(w: String) -> String
extern fn cp_is_negation(w: String) -> Bool
extern fn cp_is_neg_adverb(w: String) -> Bool
extern fn cp_irr2(surface: String) -> [String]
extern fn cp_reg_verb(w: String) -> [String]
extern fn cp_analyze_verb(surface: String) -> [String]
extern fn cp_verb_start(toks: [String], end: Int) -> Int
extern fn cp_subord_start(toks: [String], n: Int) -> Int
-287
View File
@@ -1,287 +0,0 @@
// dialogue.el SUMMON-THROUGH-SELF, native el. Port of dialogue.py's core.
//
// THE WHOLE DIALOGUE IS ONE OPERATION. A fact is never merely *fetched*: the
// query is PROJECTED into the engram's self + memory geometry, LANDS in a region,
// and the reply is READ OUT / the region MATERIALIZED from wherever it landed.
//
// project(query) -> land on a region -> read out from that region
//
// lands in the SELF region -> grounded identity/presence, read out of
// the real self nodes (self_region.el)
// lands on a memory NEIGHBORHOOD -> MATERIALIZE it: walk the neighborhood
// (engram_neighbors_json) and read out the
// region's connected members
// lands nowhere close -> HONEST ABSENCE (an empty region, not a
// fabricated answer, not an error)
//
// CRITICAL INVARIANTS (enforced structurally, not by convention):
// * ONE operation there is NO intent classifier and NO separate
// fact-retrieval branch. Identity is nearest-region proximity, not a switch.
// * MATERIALIZE by walking the neighborhood, never by fetching top-props.
// * HONEST ABSENCE when the region is thin.
// * NEGATION is SACRED: the readout is the stored prose VERBATIM, so a negated
// memory stays negated we never paraphrase a polarity away.
// * NO ECHO: the old "I noted that X. That relates to Y." template is gone.
// The summon path materializes or honestly declines it never echoes.
// * DIRECTIVE OVERRIDE: a meta-directive ("answer in English") overrides the
// reply language while the content language is still auto-detected.
//
// Depends on: comprehend (parse_spec_lang, cp_tokenize), multilingual (ml_detect,
// ml_tr, ml_term), propositions (prop_split_sentences), self_region
// (sr_available, sr_readout), the engram + json runtime builtins.
// directive override
// Return [target_lang, content]. target_lang is "" when no directive is present.
// A directive names an output language; we strip it and keep the remaining text
// as the content (whose OWN language is still auto-detected downstream).
fn dlg_dir_hit(low: String, phrase: String) -> Bool {
return str_contains(low, phrase)
}
fn dlg_parse_directive(text: String) -> [String] {
let low: String = str_to_lower(text)
let lang: String = ""
let phrase: String = ""
// English target
if dlg_dir_hit(low, "in english") { let lang = "en"; let phrase = "in english" }
if dlg_dir_hit(low, "em inglês") { let lang = "en"; let phrase = "em inglês" }
if dlg_dir_hit(low, "em ingles") { let lang = "en"; let phrase = "em ingles" }
if dlg_dir_hit(low, "en inglés") { let lang = "en"; let phrase = "en inglés" }
// Portuguese target
if dlg_dir_hit(low, "in portuguese") { let lang = "pt"; let phrase = "in portuguese" }
if dlg_dir_hit(low, "em português") { let lang = "pt"; let phrase = "em português" }
// Spanish target
if dlg_dir_hit(low, "in spanish") { let lang = "es"; let phrase = "in spanish" }
if dlg_dir_hit(low, "en español") { let lang = "es"; let phrase = "en español" }
// Italian target
if dlg_dir_hit(low, "in italian") { let lang = "it"; let phrase = "in italian" }
let content: String = text
if !str_eq(phrase, "") {
// strip the directive phrase (and a common "answer"/"responda" lead-in),
// leaving the real question as content.
let idx: Int = str_index_of(low, phrase)
if idx >= 0 {
let before: String = str_slice(text, 0, idx)
let after: String = str_slice(text, idx + str_len(phrase), str_len(text))
let content = str_trim(before + " " + after)
}
// trim a leading "answer"/"responda"/"reply" and stray colon/comma.
let cl: String = str_to_lower(content)
if str_starts_with(cl, "answer") { let content = str_trim(str_slice(content, 6, str_len(content))) }
if str_starts_with(cl, "responda") { let content = str_trim(str_slice(content, 8, str_len(content))) }
if str_starts_with(cl, "reply") { let content = str_trim(str_slice(content, 5, str_len(content))) }
if str_starts_with(content, ":") { let content = str_trim(str_slice(content, 1, str_len(content))) }
if str_starts_with(content, ",") { let content = str_trim(str_slice(content, 1, str_len(content))) }
}
let r: [String] = native_list_empty()
let r = native_list_append(r, lang)
let r = native_list_append(r, content)
return r
}
// identity landing (a region proximity, not a classifier switch)
// The query lands in the SELF region when it takes an identity/presence shape.
// Cross-lingual forms are included because the engram's lexical probe is
// English-leaning. This is the SELF attractor of the single operation.
fn dlg_is_identity(content: String) -> Bool {
let low: String = str_to_lower(str_trim(content))
if str_contains(low, "who are you") { return true }
if str_contains(low, "what are you") { return true }
if str_contains(low, "who i am") { return true }
if str_contains(low, "your name") { return true }
if str_contains(low, "about yourself") { return true }
if str_contains(low, "are you conscious") { return true }
if str_contains(low, "are you there") { return true }
// cross-lingual identity question-forms
if str_contains(low, "quem é você") { return true }
if str_contains(low, "quem es voce") { return true }
if str_contains(low, "quién eres") { return true }
if str_contains(low, "quien eres") { return true }
if str_contains(low, "chi sei") { return true }
if str_contains(low, "qui es-tu") { return true }
if str_contains(low, "wer bist du") { return true }
return false
}
// readout helpers
fn dlg_first_sentence(content: String) -> String {
let sents: [String] = prop_split_sentences(content)
let n: Int = native_list_len(sents)
let i: Int = 0
while i < n {
let s: String = str_trim(native_list_get(sents, i))
// drop a leading markdown heading marker for a clean read-out line
if str_starts_with(s, "# ") { let s = str_trim(str_slice(s, 2, str_len(s))) }
if str_len(s) > 0 { return s }
let i = i + 1
}
return str_trim(content)
}
// strip trailing/leading punctuation from a token.
fn dlg_clean_tok(w: String) -> String {
let s: String = str_trim(w)
let s = str_strip_suffix(s, ".")
let s = str_strip_suffix(s, ",")
let s = str_strip_suffix(s, "?")
let s = str_strip_suffix(s, "!")
let s = str_strip_suffix(s, ":")
let s = str_strip_suffix(s, ";")
return str_trim(s)
}
// closed-class across the supported languages (union) a word we must NOT treat
// as a retrieval topic. Also drops the meta verbs of a request ("tell", "prove",
// "show") so the TOPIC, not the speech act, is what projects into memory.
fn dlg_is_stop(w: String) -> Bool {
if ml_stop_en(w) { return true }
if ml_stop_es(w) { return true }
if ml_stop_pt(w) { return true }
if ml_stop_it(w) { return true }
if str_eq(w, "tell") { return true }
if str_eq(w, "show") { return true }
if str_eq(w, "about") { return true }
if str_eq(w, "sobre") { return true }
if str_eq(w, "acerca") { return true }
return false
}
// The CONTENT TERMS the query projects into memory: content words only, cleaned,
// cross-lingually mapped to the engram's English vocabulary, 3 chars. This is
// the geometry probe the speech-act verbs and function words are stripped so a
// PP topic ("tell me ABOUT Lisbon") projects on "lisbon", not "tell"/"me".
fn dlg_content_terms(content: String, lang: String) -> [String] {
let toks: [String] = cp_tokenize(content)
let n: Int = native_list_len(toks)
let out: [String] = native_list_empty()
let i: Int = 0
while i < n {
let w: String = str_to_lower(dlg_clean_tok(native_list_get(toks, i)))
if str_len(w) >= 3 {
if !dlg_is_stop(w) {
let out = native_list_append(out, ml_term(w, lang))
}
}
let i = i + 1
}
return out
}
// Does this landed node lexically overlap the query's content terms? This is the
// RELEVANCE FLOOR: activation always returns the store's most salient nodes, so
// without this a query about nothing would "land" on the self/top node. A node
// that shares no content term with the query is "nowhere close" -> honest absence.
fn dlg_node_matches(node: String, terms: [String]) -> Bool {
let hay: String = str_to_lower(json_get_string(node, "content") + " " + json_get_string(node, "label"))
let n: Int = native_list_len(terms)
let i: Int = 0
while i < n {
let t: String = native_list_get(terms, i)
if str_len(t) >= 3 {
if str_contains(hay, t) { return true }
}
let i = i + 1
}
return false
}
// MATERIALIZE the landed region: read out the landed fact, then WALK the
// neighborhood and read out its connected members (real edges, not top-props).
fn dlg_materialize(top_node: String, reply_lang: String) -> String {
let id: String = json_get_string(top_node, "id")
let content: String = json_get_string(top_node, "content")
let lead: String = dlg_first_sentence(content)
let nb: String = engram_neighbors_json(id, 2, "both")
let m: Int = json_array_len(nb)
let parts: [String] = native_list_empty()
let parts = native_list_append(parts, lead)
let added: Int = 0
let i: Int = 0
while i < m {
if added < 3 {
let rec: String = json_array_get(nb, i)
let node: String = json_get_raw(rec, "node")
let nc: String = json_get_string(node, "content")
if !str_eq(nc, "") {
let sent: String = dlg_first_sentence(nc)
if !str_eq(sent, "") {
let parts = native_list_append(parts, sent)
let added = added + 1
}
}
}
let i = i + 1
}
// The readout is the region's OWN prose, verbatim negation SACRED, no echo.
return str_join(parts, " ")
}
// THE single operation
fn dlg_respond(text: String) -> String {
// directive override: reply language may differ from content language.
let dir: [String] = dlg_parse_directive(text)
let target_lang: String = native_list_get(dir, 0)
let content: String = native_list_get(dir, 1)
let content_lang: String = ml_detect(content)
let reply_lang: String = content_lang
if !str_eq(target_lang, "") { let reply_lang = target_lang }
// comprehend the content (SACRED polarity carried in the spec).
let spec: [String] = parse_spec_lang(content, content_lang)
// PROJECT + LAND: SELF region
// Identity/presence shape lands in the self region; read out the REAL self
// nodes (self_region.el), never a template. Same single operation this is
// just the self attractor winning the landing.
if dlg_is_identity(content) {
if sr_available() {
// read out the REAL self nodes when replying in their own language
// (the soul's prose is English); for another reply language we cannot
// translate real content without an LLM, so we answer with the
// localized SACRED identity anchor honest, in-language, no fabrication.
if str_eq(reply_lang, "en") { return sr_readout("en") }
return ml_tr("identity", reply_lang)
}
// self region thin honest localized identity (logged fallback shape).
return ml_tr("identity", reply_lang)
}
// PROJECT into MEMORY geometry
let terms: [String] = dlg_content_terms(content, content_lang)
let qterm: String = str_join(terms, " ")
let act: String = engram_activate_json(qterm, 12)
let n: Int = json_array_len(act)
// LAND: the highest-activation node that ACTUALLY overlaps the query's
// content terms (the relevance floor). Activation always returns the most
// salient nodes, so we walk the ranked list and take the first that is
// genuinely "close"; if none is, the query landed nowhere. ───────────────
let landing: String = ""
let i: Int = 0
while i < n {
if str_eq(landing, "") {
let rec: String = json_array_get(act, i)
let node: String = json_get_raw(rec, "node")
if dlg_node_matches(node, terms) {
let landing = node
}
}
let i = i + 1
}
// HONEST ABSENCE: nothing close an empty region, not a fabricated answer,
// not an "I noted that" echo.
if str_eq(landing, "") {
return ml_tr("no_memory", reply_lang)
}
// MATERIALIZE the landing by WALKING its neighborhood.
return dlg_materialize(landing, reply_lang)
}
-13
View File
@@ -63,9 +63,6 @@ import "morphology-cop.el"
import "grammar.el"
import "realizer.el"
import "semantics.el"
// Comprehension front-end (input half: text meaning-spec)
import "comprehend.el"
//
// Entry points:
//
@@ -120,9 +117,6 @@ fn build_form_from_json(semantic_form_json: String, lang_code: String) -> [Strin
let location: String = sem_get(semantic_form_json, "location")
let tense: String = sem_get(semantic_form_json, "tense")
let aspect: String = sem_get(semantic_form_json, "aspect")
let polarity: String = sem_get(semantic_form_json, "polarity")
let neg_word: String = sem_get(semantic_form_json, "neg_word")
let iobj: String = sem_get(semantic_form_json, "iobj")
let form: [String] = native_list_empty()
let form = native_list_append(form, "intent")
@@ -133,19 +127,12 @@ fn build_form_from_json(semantic_form_json: String, lang_code: String) -> [Strin
let form = native_list_append(form, predicate)
let form = native_list_append(form, "patient")
let form = native_list_append(form, patient)
let form = native_list_append(form, "iobj")
let form = native_list_append(form, iobj)
let form = native_list_append(form, "location")
let form = native_list_append(form, location)
let form = native_list_append(form, "tense")
let form = native_list_append(form, tense)
let form = native_list_append(form, "aspect")
let form = native_list_append(form, aspect)
// SACRED: polarity crosses the JSON boundary and is never inferred away.
let form = native_list_append(form, "polarity")
let form = native_list_append(form, polarity)
let form = native_list_append(form, "neg_word")
let form = native_list_append(form, neg_word)
let form = native_list_append(form, "lang")
let form = native_list_append(form, lang_code)
-72
View File
@@ -1,72 +0,0 @@
;;; lang_profile_ca.el — Catalan language profile for ELP.
;;; Mirrors lang_profile_it / _es / _pt; keys the realizer's construction switches.
;;; Catalan is the CLOSEST Romance sibling to the shared engine (~85% conceptual
;;; reuse). The deltas: PRONOMS FEBLES with four position allomorphs, l'-elision,
;;; del/al/pel contractions, the periphrastic preterite (vaig+INF), and NO
;;; essere/avere split (perfect aux is always HAVER; ser/estar is only the copula).
(lang_profile_ca
(language "Catalan")
(iso639 "ca")
(family "Romance")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop yes) ; null subjects default; overt pronoun = emphatic
(obligatory-subject no)
(grammatical-gender yes) ; m/f; full NP agreement (art + adj + participle)
(do-support no)
(subject-aux-inversion no) ; yes/no Q = declarative order + '?'; no inversion
(article-selection "el/la/l'/els/les ; un/una/uns/unes") ; l'-ELISION:
; el/la -> l' before vowel or (silent) h, glued to
; the next word (l'home, l'illa); de -> d' before vowel
(article-drives-contraction yes) ; article choice feeds prep+article contraction
(adjective-position "postnominal-default + small prenominal class") ; bo/bon,
; mal, gran, nou, vell, primer, molt... prenominal
(question-punct plain) ; ? and ! only (no inverted ¿ ¡)
;; ── MANDATORY prep+article contractions ────────────────────────────────
(contractions ((de el del) (de els dels)
(a el al) (a els als)
(per el pel) (per els pels)))
(contraction-mandatory yes) ; *de el -> del obligatory
(contraction-blocked-before-elision yes) ; de l'home / a l'home (NO *del home)
;; ── clitic system: PRONOMS FEBLES (the headline delta) ──────────────────
(clitics yes)
(clitic-allomorphy four-position) ; per pronoun, form varies by position+onset:
; reinforced (em, et, el) proclitic before a consonant
; elided (m', t', l', n') proclitic before a vowel/h
; full (-me, -lo, -li) enclitic after a consonant/-r
; reduced ('m, 't, 'l, 'ns) enclitic after a vowel
(clitic-placement ((finite proclitic) ; el veig, no m'ho dóna
(imperative-affirmative enclitic) ; dóna'm, digues-me
(imperative-negative present-subjunctive) ; no parlis (delta)
(infinitive enclitic) ; ajudar-me, veure'l
(gerund enclitic))) ; fent-ho
(clitic-combination ((me el "me'l") (te el "te'l") (se el "se'l")
(me la "me la") (me en "me'n")
(li el "l'hi") (li en "n'hi"))) ; dative+accusative clusters
(clitic-particles (hi en ho)) ; locative hi, partitive/genitive en, neuter ho
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "person+number (6-way)")
(tenses (present imperfet preterit-simple perifrastic-preterit futur
condicional subjuntiu-present subjuntiu-imperfet imperatiu))
(periphrastic-preterite "vaig/vas/va/vam/vau/van + INFINITIVE") ; << hallmark CA
; (vaig cantar = 'I sang'); coexists w/ synthetic pret.
(compound-past "pretèrit perfet = haver(present) + participle")
(perfect-aux "HAVER only") ; << NO essere/avere split (simpler than IT)
(participle-agreement ((haver preceding-acc-clitic))) ; les he vistes; else invariable
(progressive-aux "estar + gerundi")
(copula "ser / estar") ; ser: identity/essential/origin; estar:
; location + transient state (estic cansat, és a casa)
(passive-aux "ser (+ per-agent)")
(future inflectional) ; cantaré, serà
(comparative "més/menys ADJ que")
;; ── SACRED safety bar (shared with es/pt/it/en) ────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
(negation "no (preverbal) + optional 'pas' + concord") ; no...res/
; ningú/mai/cap/gens/enlloc
(negative-concord yes) ; preverbal negative subject (ningú) keeps 'no'
(neg-reinforcer pas)) ; optional (no ho faré pas)
-41
View File
@@ -1,41 +0,0 @@
;;; lang_profile_de.el — German language profile for ELP.
;;; Mirrors lang_profile_en / lang_profile_es. Keys the realizer's construction
;;; switches. German is the largest Germanic delta from the EN engine: V2 word
;;; order, four morphological cases, and separable-prefix verbs.
(lang_profile_de
(language "German")
(iso639 "de")
(family "Germanic")
(neighbor-base "en") ; realized by extending the English (Germanic) engine
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop no) ; obligatory subject in finite clauses
(obligatory-subject yes)
(grammatical-gender (m f n)) ; three genders; drives article + adj declension
(case-system (nom acc dat gen)) ; four cases on articles/adjs/nouns
(word-order V2) ; finite verb 2nd in main clause
(subordinate-order verb-final) ; "..., dass er den Hund SIEHT."
(separable-verbs yes) ; aufstehen -> "steht ... auf"; ppart "aufgestanden"
(do-support no) ; German negates/questions the finite verb directly
(subject-verb-inversion yes) ; yes/no Q fronts finite verb; wh-Q fills Vorfeld
(article-selection "der/die/das + ein/kein") ; declined by case x gender x number
(adjective-position prenominal)
(adjective-declension (strong weak mixed)) ; chosen by the determiner type
(noun-capitalization yes)
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "person-and-number") ; full present/past paradigm
(auxiliary-order (modal tense-aux perfect passive main))
(perfect-aux (haben sein)) ; sein for intransitive motion/change verbs
(passive-aux "werden")
(future "werden + infinitive")
(comparative "synthetic (-er / -st, with umlaut)")
;; ── negation ───────────────────────────────────────────────────────────
(negation-markers (nicht kein)) ; kein- negates an indefinite NP; nicht else
(negation-faithful yes) ; SACRED: polarity never dropped/inverted -> FLAG
;; ── lexicon provenance ─────────────────────────────────────────────────
(lexicon-source "UniMorph deu (primary) + kaikki.org German (gender override)")
(lexicon-license "CC-BY-SA 3.0 / GFDL"))
-41
View File
@@ -1,41 +0,0 @@
;;; lang_profile_en.el — English language profile for ELP.
;;; Mirrors lang_profile_es / lang_profile_pt; keys the realizer's construction
;;; switches. English is typologically distinct from the Romance builds, so the
;;; flags differ where the grammar differs.
(lang_profile_en
(language "English")
(iso639 "en")
(family "Germanic")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop no) ; OBLIGATORY subjects — missing subject is FLAGGED
(obligatory-subject yes)
(grammatical-gender no) ; natural gender only (he/she/it), no NP agreement
(do-support yes) ; negation & questions of lexical verbs insert do/does/did
(subject-aux-inversion yes) ; yes/no + non-subject wh questions invert the operator
(article-selection "a/an/the") ; a/an resolved PHONOLOGICALLY (an hour, a university)
(adjective-position prenominal) ; attributive adjectives precede the noun; invariant
(has-tag-questions yes) ; "...doesn't he?" — operator + reversed polarity
(has-there-existential yes) ; "there is/are/have been ..."
(possessive-clitic "'s") ; saxon genitive; plural in -s -> bare apostrophe
(question-punct plain) ; ? and ! only (no inverted marks)
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "3sg-present-only") ; only 3sg present -s (+ suppletive be)
(auxiliary-order (modal perfect progressive passive main))
(perfect-aux "have") ; have + past participle
(progressive-aux "be") ; be + present participle
(passive-aux "be") ; be + past participle (+ by-agent)
(future "will + base") ; no inflectional future
(comparative "synthetic-or-periphrastic") ; -er/-est vs more/most by syllables
;; ── SACRED safety bar (shared with es/pt) ──────────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
;; ── DIALECT overlay (post-realization, one core -> US/UK/AU) ────────────
(dialect US) ; default; profile field switches the overlay
(dialects (US UK AU))
(dialect-canonical US) ; core is authored in US orthography
(dialect-overlay "dialect_en.to_dialect") ; orthography + lexis + grammar prefs
(dialect-covers (spelling lexis collective-agreement gotten/got)))
-45
View File
@@ -1,45 +0,0 @@
;;; lang_profile_es.el — Spanish language profile for ELP.
;;; Keys the realizer's construction switches. Mirrors lang_profile_en / _pt.
(lang_profile_es
(language "Spanish")
(iso639 "es")
(family "Romance")
;; -- core typology flags -------------------------------------------------
(pro-drop yes) ; subjects routinely dropped; agreement carries person
(obligatory-subject no)
(grammatical-gender yes) ; m/f on every noun; article+adjective AGREE
(gender-source lexicon); REAL per-noun gender from UniMorph — NOT a heuristic
(do-support no)
(subject-aux-inversion no) ; questions by intonation/punctuation, not inversion
(question-strategy intonation)
(article-selection "el/la/los/las un/una/unos/unas")
(stressed-a-rule yes) ; fem sg noun in stressed a-/ha- takes el/un (el agua)
(adjective-position postnominal) ; default post; a few prenominal + apocope
(adjective-agreement "gender+number")
(question-punct inverted) ; opening ¿ ¡ required
;; -- MANDATORY CONTRACTIONS (coordinator quality bar) --------------------
(contractions ((de el "del") (a el "al")))
(contraction-mandatory yes) ; 'de el'/'a el' MUST surface as del/al
;; -- verb / aspect system ------------------------------------------------
(verb-classes (ar er ir))
(tenses (present preterite imperfect future conditional))
(moods (ind sbjv imp))
(finite-agreement "person+number (6 slots)")
(perfect-aux "haber") ; haber + past participle (invariant -o)
(progressive-aux "estar") ; estar + gerund
(passive-aux "ser") ; ser + participle (agrees) + por-agent
(copula-split "ser/estar") ; permanent vs stage-level
(future "infinitive + é/ás/á/emos/éis/án")
;; -- clitics / government ------------------------------------------------
(object-clitics yes) ; me te lo la le nos os los las; proclisis/enclisis
(clitic-order "se II I III (le+lo -> se lo)")
(enclisis "imperative/infinitive/gerund + accent repair (dá+me+lo->dámelo)")
(verb-prep-government yes) ; verbs select prep (protestar+contra, escapar+de)
;; -- SACRED safety bar (shared with en/pt) -------------------------------
(negation-faithful yes)) ; polarity never dropped/inverted; unplaceable -> FLAG
-74
View File
@@ -1,74 +0,0 @@
;;; lang_profile_fr.el — French language profile for ELP.
;;; Mirrors lang_profile_it / lang_profile_es; keys the realizer's construction
;;; switches. French is a Romance sibling (~54% of the realizer code and the whole
;;; clause-engine architecture reused), but carries the family's biggest surface
;;; deltas: NOT pro-drop, DISCONTINUOUS negation, and an orthography/phonology
;;; mismatch (elision, liaison) that makes exact-match genuinely hard.
(lang_profile_fr
(language "French")
(iso639 "fr")
(family "Romance")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop no) ; << French-specific: subject clitic OBLIGATORY
(obligatory-subject yes) ; je/tu/il/elle/nous/vous/ils/elles always overt
(grammatical-gender yes) ; m/f; full NP agreement (art + adj + participle)
(do-support no)
(subject-aux-inversion optional) ; est-ce que (default) OR clitic inversion (vas-tu)
(article-selection "le/la/l'/les ; un/une/des ; PARTITIVE du/de la/de l'/des")
(article-drives-contraction yes) ; à+le=au, de+le=du feed off article choice
(adjective-position "postnominal-default + prenominal-BAGS") ; beau/bon/grand/
; petit/jeune/vieux/nouveau + ordinals prenominal
; (beau->bel, nouveau->nouvel, vieux->vieil / vowel)
(question-punct "space-before") ; French typography: ' ?' ' !' (no ¿¡)
;; ── elision (orthography/phonology mismatch — French-specific) ──────────
(elision ((le l') (la l') (je j') (ne n') (de d') (que qu')
(me m') (te t') (se s') (ce c'))) ; before vowel / h-muet
(elision-h-muet yes) ; l'homme, l'hôpital (h-aspiré exception list kept)
(liaison noted-not-modeled) ; phonological, not written in surface
;; ── MANDATORY prep+article contractions ────────────────────────────────
(contractions ((à le au) (à les aux) (de le du) (de les des)))
(contraction-mandatory yes) ; *à le -> au obligatory; à la / à l' uncontracted
(partitive ((m-sg du) (f-sg "de la") (vowel "de l'") (pl des)))
(partitive-under-neg "de") ; << gap in current build: 'ne … pas de pain'
;; ── clitic system ──────────────────────────────────────────────────────
(clitics yes)
(clitic-order (me te se nous vous | le la les | lui leur | y | en))
(clitic-placement ((finite proclitic) ; je le lui donne
(imperative-affirmative enclitic-hyphen) ; donne-le-moi
(imperative-negative "ne+proclitic+verb+pas") ; ne le donne pas
(infinitive enclitic))) ; PARTIAL: clitic-climbing
; onto infinitive under modal
(clitic-imperative-shift ((me moi) (te toi))) ; final me/te -> moi/toi (donne-moi)
(clitic-particles (y en)) ; locative y, partitive/genitive en
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "person+number (written; many homophones)")
(tenses (présent imparfait passé-simple futur conditionnel
subjonctif-présent subjonctif-imparfait impératif))
(compound-past "passé-composé = aux(present) + participe passé")
(perfect-aux "être/avoir (LEXICAL selection)") ; << French-specific
(etre-aux-class "intransitive motion/change (aller venir arriver partir
entrer sortir monter descendre naître mourir rester
tomber retourner passer devenir revenir rentrer) + ALL
pronominal verbs")
(participle-agreement ((être subject) ; elle est allée / elles venues
(avoir preceding-direct-object))) ; je les ai vus
(progressive "être en train de + infinitif") ; no dedicated aux
(copula "être (single; no ser/estar, no essere/stare)")
(passive-aux "être (+ par-agent)")
(future inflectional) ; parlera, sera
(comparative "plus/moins ADJ que")
(superlative "le/la plus ADJ (de …)") ; PARTIAL word-order in build
;; ── SACRED safety bar (shared with es/pt/it/en) ────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
(negation "DISCONTINUOUS: ne (preverbal) … pas/jamais/rien/personne/
plus/guère/que (postverbal)") ; << biggest structural delta
(negation-ne-elides yes) ; ne -> n' before vowel (n'ai pas vu)
(negation-passe-composé "ne + aux + pas + participe") ; n'ai pas vu
(negative-concord partial)) ; personne/rien as arguments post-participle
-70
View File
@@ -1,70 +0,0 @@
;;; lang_profile_it.el — Italian language profile for ELP.
;;; Mirrors lang_profile_es / lang_profile_pt; keys the realizer's construction
;;; switches. Italian is a Romance sibling, so ~85% of the flags match ES/PT; the
;;; essere/avere auxiliary split and phonological article selection are the deltas.
(lang_profile_it
(language "Italian")
(iso639 "it")
(family "Romance")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop yes) ; null subjects default; overt pronoun = emphatic
(obligatory-subject no)
(grammatical-gender yes) ; m/f; full NP agreement (art + adj + participle)
(do-support no)
(subject-aux-inversion no) ; yes/no Q = declarative order + '?'; no inversion
(article-selection "il/lo/l'/i/gli + la/l'/le ; un/uno/un'/una") ; PHONOLOGICAL:
; lo/gli/uno before s+cons, z, gn, ps, pn, x, y, i+V;
; l'/un' before a vowel (elision, glued to next word)
(article-drives-contraction yes) ; article choice feeds the prep+art contraction
(adjective-position "postnominal-default + prenominal-class") ; bello/buono/grande
; /nuovo/vecchio/primo... prenominal (with apocope)
(question-punct plain) ; ? and ! only (no inverted ¿ ¡)
;; ── MANDATORY prep+article contractions ────────────────────────────────
(contractions ((di il del) (di lo dello) (di la della) (di i dei)
(di gli degli) (di le delle) (di l' dell')
(a il al) (a lo allo) (a la alla) (a i ai) (a gli agli)
(a le alle) (a l' all')
(da il dal) (da la dalla) (da gli dagli) (da l' dall')
(in il nel) (in la nella) (in gli negli) (in l' nell')
(su il sul) (su la sulla) (su gli sugli) (su l' sull')))
(contraction-mandatory yes) ; *di il -> del is obligatory, never uncontracted
(prep-no-contract (per tra fra)) ; per la strada (NOT *perla)
;; ── clitic system ──────────────────────────────────────────────────────
(clitics yes)
(clitic-placement ((finite proclitic) ; lo vedo, non me lo dà
(imperative-affirmative enclitic) ; dammelo, guardalo
(imperative-negative-tu non+infinitive) ; non parlare / non lo fare
(infinitive enclitic) ; vederlo, aiutarmi (drop -e)
(gerund enclitic))) ; dandolo
(clitic-combination ((mi lo "me lo") (ti lo "te lo") (ci lo "ce lo")
(vi lo "ve lo") (si lo "se lo")
(gli lo "glielo") (le lo "glielo"))) ; glielo = ONE word
(clitic-particles (ci ne)) ; locative ci, partitive ne
(raddoppiamento (da fa di va sta)) ; monosyllabic imper double clitic: dammelo
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "person+number (6-way)")
(tenses (presente imperfetto passato-remoto futuro condizionale
congiuntivo-presente congiuntivo-imperfetto imperativo))
(compound-past "passato-prossimo = aux(present) + participle")
(perfect-aux "essere/avere (LEXICAL selection)") ; << Italian-specific
(essere-aux-class unaccusative) ; motion/change-of-state/copular/pronominal
; (andare venire nascere morire diventare piacere
; + ALL reflexives) -> essere
(participle-agreement ((essere subject) ; è andata / sono arrivati
(avere preceding-acc-clitic))) ; li ho visti
(progressive-aux "stare + gerundio") ; sto parlando
(copula "essere (default) / stare (state: sto bene)")
(passive-aux "essere / venire (+ da-agent)")
(future inflectional) ; parlerò, sarà
(comparative "più/meno ADJ di")
;; ── SACRED safety bar (shared with es/pt/en) ───────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
(negation "non (preverbal) + concord") ; non...niente/nessuno/mai/più
(negative-concord yes) ; preverbal negative word (nessuno/niente) suppresses non
(neg-adverb-position between-aux-and-participle)) ; non ho MAI visto
-30
View File
@@ -1,30 +0,0 @@
;;; lang_profile_la.el — Latin language profile for ELP.
;;; Keys the realizer's construction switches. Companion to morphology-la.el.
(lang_profile_la
(language "Latin")
(iso639 "la")
(family "Italic")
;; -- core typology flags -------------------------------------------------
(pro-drop yes) ; person carried by verb ending; subjects dropped
(obligatory-subject no)
(grammatical-gender yes) ; m/f/n; adjective AGREES in case+gender+number
(gender-source lexicon) ; REAL per-noun gender from UniMorph lat
(articles none) ; Latin has no articles
(case-system yes) ; NOM GEN DAT ACC ABL VOC (+ rare LOC)
(cases (nom gen dat acc abl voc))
(word-order "SOV (default; free order, case-marked)")
(adjective-position "either (case agreement carries the link)")
(adjective-agreement "case+gender+number")
;; -- verb / aspect system ------------------------------------------------
(verb-classes (1 2 3 3io 4)) ; four conjugations + i-stem 3rd
(tenses (present imperfect future perfect pluperfect futureperfect))
(moods (indicative subjunctive imperative infinitive))
(voices (active passive))
(finite-agreement "person+number (6 slots)")
(citation "principal parts: pres-1sg / pres-inf / perf-participle")
;; -- SACRED safety bar ---------------------------------------------------
(negation-faithful yes)) ; polarity never dropped/inverted
-40
View File
@@ -1,40 +0,0 @@
;;; lang_profile_pt.el — Portuguese language profile for ELP.
;;; Keys the realizer's construction switches. Mirrors lang_profile_es.
(lang_profile_pt
(language "Portuguese")
(iso639 "pt")
(family "Romance")
;; -- core typology flags -------------------------------------------------
(pro-drop yes) ; subjects routinely dropped; agreement carries person
(obligatory-subject no)
(grammatical-gender yes) ; m/f on every noun; article+adjective AGREE
(gender-source lexicon) ; REAL per-noun gender from UniMorph por / kaikki
(do-support no)
(subject-aux-inversion no)
(question-strategy intonation)
(article-selection "o/a/os/as um/uma/uns/umas")
(adjective-position postnominal)
(adjective-agreement "gender+number")
;; -- MANDATORY CONTRACTIONS (prep + article) -----------------------------
(contractions ((de o "do") (de a "da") (em o "no") (em a "na")
(a o "ao") (a a "à") (por o "pelo") (por a "pela")))
(contraction-mandatory yes)
;; -- verb / aspect system ------------------------------------------------
(verb-classes (ar er ir))
(tenses (present preterite imperfect future conditional))
(moods (ind sbjv imp))
(finite-agreement "person+number (6 slots)")
(perfect-aux "ter") ; ter + past participle
(copula-split "ser/estar")
(personal-infinitive yes) ; distinctive PT inflected infinitive
;; -- clitics / government ------------------------------------------------
(object-clitics yes) ; mesoclisis/enclisis/proclisis by context
(verb-prep-government yes)
;; -- SACRED safety bar ---------------------------------------------------
(negation-faithful yes))
-71
View File
@@ -1,71 +0,0 @@
;;; lang_profile_ro.el — Romanian language profile for ELP.
;;; Romanian is the BIG typological delta of the Romance family. The verb/clause
;;; engine and the SACRED negation contract mirror the ES/PT/IT core, but the
;;; NOMINAL system is genuinely new: a SUFFIXED definite article, preserved CASE,
;;; a NEUTER gender, and a VOCATIVE. Those flags mark where the shared engine was
;;; extended rather than reused.
(lang_profile_ro
(language "Romanian")
(iso639 "ro")
(family "Romance (Eastern / Balkan)")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop yes) ; null subjects default; overt pronoun = emphatic
(obligatory-subject no)
(grammatical-gender yes) ; m / f / NEUTER (n)
(neuter-gender yes) ; << ROMANIAN-SPECIFIC: masc-agreeing SG, fem-agreeing PL
; (un tren nou / două trenuri noi)
(do-support no)
(subject-aux-inversion no) ; yes/no Q = declarative order + '?'
(question-punct plain) ; ? and ! only
;; ── SUFFIXED DEFINITE ARTICLE (the headline engine extension) ───────────
(definite-article suffixed) ; << UNIQUE IN ROMANCE: enclitic on the noun
(definite-forms ((m/n sg "-ul / -le / -l : om->omul, câine->câinele, codru->codrul")
(f sg "-a / -ea / -ua : casă->casa, carte->cartea, stea->steaua")
(m pl "-i : oameni->oamenii")
(f/n pl "-le : case->casele, trenuri->trenurile")))
(article-host ((no-prenom-adj noun) ; omul bun
(prenom-adj adjective))) ; bunul om (adj carries the article)
(indefinite-article ((m/n "un") (f "o") (pl "niște") (gen/dat-pl "unor")))
;; ── CASE (preserved; NOM/ACC vs GEN/DAT) ────────────────────────────────
(case (nom/acc gen/dat vocative)) ; << ROMANIAN-SPECIFIC
(case-syncretism "nom=acc ; gen=dat")
(genitive-marking "gen/dat definite: -lui (m/n), -ei/-i (f), -lor (pl)")
(genitival-article ((m sg "al") (f sg "a") (m pl "ai") (f/n pl "ale"))) ; o carte a lui
(possession "definite-head + gen/dat possessor: casa băiatului")
(vocative ((m sg "-ule/-e : omule, băiete") (f sg "-o : Mario, fato")
(pl "-lor")))
;; ── verb / aspect system ────────────────────────────────────────────────
(finite-agreement "person+number (6-way)")
(tenses (prezent imperfect perfect-simplu conjunctiv-prezent
imperativ (periphrastic: perfect-compus viitor conditional)))
(compound-past "perfectul compus = a-avea-clitic + INVARIABLE participle")
(perfect-aux "a avea (am/ai/a/am/ați/au) — ONE auxiliary for ALL verbs")
(perfect-aux-split no) ; << SIMPLER than Italian: no essere/avere selection
(participle-agreement none) ; invariable in the perfect compus (agrees only as
; an adjective / in the passive)
(future "voi/vei/va/vom/veți/vor + infinitive (viitor literar)")
(conditional "aș/ai/ar/am/ați/ar + infinitive")
(subjunctive "conjunctiv: particle 'să' + subjunctive present")
(modal-complement "modal + să + subjunctive (vreau să merg, poți să ajuți)")
(copula "a fi")
(passive "a fi + participle (participle AGREES like an adjective)")
(comparative "mai / mai puțin ADJ decât")
;; ── clitic system (partial — see honest gaps) ───────────────────────────
(clitics yes)
(clitic-set ((acc te îl o ne îi le) (dat îmi îți îi ne le)
(refl te se ne se)))
(clitic-placement ((finite proclitic) ; îmi place, o văd
(perfect-compus elision) ; << m-am, l-am, i-am (PARTIAL)
(imperative-affirmative enclitic))) ; dă-mi (PARTIAL)
;; ── SACRED safety bar (shared with es/pt/it/en) ─────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
(negation "nu (single preverbal marker) + concord")
(negative-concord yes) ; nu … nimic / nimeni / niciodată / niciun
(negative-imperative "nu + INFINITIVE : nu pleca! (KNOWN GAP: uses imperative stem)"))
-1
View File
@@ -250,7 +250,6 @@ fn en_irregular_verb(base: String) -> [String] {
if str_eq(base, "cut") { let r: [String] = ["cut", "cuts", "cut", "cut", "cutting"]; return r }
if str_eq(base, "set") { let r: [String] = ["set", "sets", "set", "set", "setting"]; return r }
if str_eq(base, "hit") { let r: [String] = ["hit", "hits", "hit", "hit", "hitting"]; return r }
if str_eq(base, "fight") { let r: [String] = ["fight", "fights","fought", "fought", "fighting"]; return r }
return empty
}
-280
View File
@@ -1,280 +0,0 @@
// multilingual.el - the language layer for the native-el interlocutor.
//
// Deterministic, NO generative model (ports multilingual.py):
// 1. ml_detect(text) -> ISO code (en/es/pt/it) via stopword + diacritic score
// 2. ml_tr(key, lang) -> localized fixed phrase (SACRED per-language yes/no/decline)
// 3. ml_term(w, lang) -> PT/ES content term -> EN engram equivalent
// 4. ml_translate_pred(lemma, lang) -> EN predicate lemma -> target infinitive
//
// The Python detector count-weights stopwords and diacritics; here diacritics are
// scored by PRESENCE (str_contains) rather than codepoint counting, to stay clear
// of UTF-8 index hazards in the runtime. Faithful enough to classify typical
// queries; documented simplification. Depends on: comprehend (cp_tokenize).
// 1. language detection
fn ml_stop_en(w: String) -> Bool {
if str_eq(w, "the") { return true }
if str_eq(w, "does") { return true }
if str_eq(w, "do") { return true }
if str_eq(w, "did") { return true }
if str_eq(w, "what") { return true }
if str_eq(w, "who") { return true }
if str_eq(w, "is") { return true }
if str_eq(w, "are") { return true }
if str_eq(w, "how") { return true }
if str_eq(w, "you") { return true }
if str_eq(w, "your") { return true }
if str_eq(w, "of") { return true }
if str_eq(w, "to") { return true }
if str_eq(w, "and") { return true }
if str_eq(w, "for") { return true }
if str_eq(w, "explain") { return true }
if str_eq(w, "answer") { return true }
if str_eq(w, "memory") { return true }
if str_eq(w, "with") { return true }
if str_eq(w, "not") { return true }
if str_eq(w, "store") { return true }
return false
}
fn ml_stop_es(w: String) -> Bool {
if str_eq(w, "que") { return true }
if str_eq(w, "qué") { return true }
if str_eq(w, "una") { return true }
if str_eq(w, "usted") { return true }
if str_eq(w, "su") { return true }
if str_eq(w, "cómo") { return true }
if str_eq(w, "como") { return true }
if str_eq(w, "cuál") { return true }
if str_eq(w, "quién") { return true }
if str_eq(w, "está") { return true }
if str_eq(w, "es") { return true }
if str_eq(w, "los") { return true }
if str_eq(w, "las") { return true }
if str_eq(w, "del") { return true }
if str_eq(w, "al") { return true }
if str_eq(w, "explica") { return true }
if str_eq(w, "explique") { return true }
if str_eq(w, "forma") { return true }
if str_eq(w, "con") { return true }
if str_eq(w, "memoria") { return true }
if str_eq(w, "responde") { return true }
return false
}
fn ml_stop_pt(w: String) -> Bool {
if str_eq(w, "que") { return true }
if str_eq(w, "uma") { return true }
if str_eq(w, "você") { return true }
if str_eq(w, "sua") { return true }
if str_eq(w, "seu") { return true }
if str_eq(w, "como") { return true }
if str_eq(w, "memória") { return true }
if str_eq(w, "isso") { return true }
if str_eq(w, "os") { return true }
if str_eq(w, "as") { return true }
if str_eq(w, "da") { return true }
if str_eq(w, "do") { return true }
if str_eq(w, "na") { return true }
if str_eq(w, "no") { return true }
if str_eq(w, "explica") { return true }
if str_eq(w, "forma") { return true }
if str_eq(w, "é") { return true }
if str_eq(w, "está") { return true }
if str_eq(w, "com") { return true }
if str_eq(w, "responda") { return true }
return false
}
fn ml_stop_it(w: String) -> Bool {
if str_eq(w, "che") { return true }
if str_eq(w, "una") { return true }
if str_eq(w, "come") { return true }
if str_eq(w, "della") { return true }
if str_eq(w, "gli") { return true }
if str_eq(w, "è") { return true }
if str_eq(w, "sono") { return true }
if str_eq(w, "questo") { return true }
if str_eq(w, "nel") { return true }
if str_eq(w, "di") { return true }
if str_eq(w, "il") { return true }
if str_eq(w, "cosa") { return true }
if str_eq(w, "per") { return true }
if str_eq(w, "memoria") { return true }
if str_eq(w, "spiega") { return true }
if str_eq(w, "rispondi") { return true }
return false
}
// diacritic PRESENCE score (weight 3 each; hard overrides weight 8).
fn ml_dia_score(low: String, lang: String) -> Int {
let s: Int = 0
if str_eq(lang, "pt") {
if str_contains(low, "ã") { let s = s + 3 }
if str_contains(low, "õ") { let s = s + 3 }
if str_contains(low, "ç") { let s = s + 3 }
if str_contains(low, "ê") { let s = s + 3 }
if str_contains(low, "á") { let s = s + 3 }
// hard PT markers (ã/õ almost never appear outside PT)
if str_contains(low, "ã") { let s = s + 8 }
if str_contains(low, "õ") { let s = s + 8 }
}
if str_eq(lang, "es") {
if str_contains(low, "ñ") { let s = s + 3 }
if str_contains(low, "¿") { let s = s + 3 }
if str_contains(low, "¡") { let s = s + 3 }
if str_contains(low, "á") { let s = s + 3 }
if str_contains(low, "é") { let s = s + 3 }
// hard ES markers
if str_contains(low, "ñ") { let s = s + 8 }
if str_contains(low, "¿") { let s = s + 8 }
if str_contains(low, "¡") { let s = s + 8 }
}
if str_eq(lang, "it") {
if str_contains(low, "è") { let s = s + 3 }
if str_contains(low, "ì") { let s = s + 3 }
if str_contains(low, "ò") { let s = s + 3 }
}
return s
}
fn ml_stop_score(toks: [String], lang: String) -> Int {
let n: Int = native_list_len(toks)
let s: Int = 0
let i: Int = 0
while i < n {
let w: String = native_list_get(toks, i)
if str_eq(lang, "en") { if ml_stop_en(w) { let s = s + 2 } }
if str_eq(lang, "es") { if ml_stop_es(w) { let s = s + 2 } }
if str_eq(lang, "pt") { if ml_stop_pt(w) { let s = s + 2 } }
if str_eq(lang, "it") { if ml_stop_it(w) { let s = s + 2 } }
let i = i + 1
}
return s
}
fn ml_detect(text: String) -> String {
if str_eq(text, "") { return "en" }
let low: String = str_to_lower(text)
let toks: [String] = cp_tokenize(text)
// NOTE: el's overloaded `+` mis-compiles two chained function-call Int operands
// as string concat (documented in comprehend_gate.el). Bind each call to an Int
// var and add vars one at a time so the addition stays integer.
let en: Int = ml_stop_score(toks, "en")
let es_s: Int = ml_stop_score(toks, "es")
let es_d: Int = ml_dia_score(low, "es")
let es: Int = es_s + es_d
let pt_s: Int = ml_stop_score(toks, "pt")
let pt_d: Int = ml_dia_score(low, "pt")
let pt: Int = pt_s + pt_d
let it_s: Int = ml_stop_score(toks, "it")
let it_d: Int = ml_dia_score(low, "it")
let it: Int = it_s + it_d
let best: String = "en"
let bs: Int = en
if es > bs { let best = "es"; let bs = es }
if pt > bs { let best = "pt"; let bs = pt }
if it > bs { let best = "it"; let bs = it }
// weak signal -> honest fallback to English
if bs < 3 { return "en" }
return best
}
// 2. localized fixed phrases (SACRED per-language decline/yes/no)
fn ml_tr(key: String, lang: String) -> String {
if str_eq(key, "no_memory") {
if str_eq(lang, "pt") { return "Não tenho isso na minha memória." }
if str_eq(lang, "es") { return "No tengo eso en mi memoria." }
if str_eq(lang, "it") { return "Non ho quello nella mia memoria." }
return "I don't have that in my memory."
}
if str_eq(key, "parse_fail") {
if str_eq(lang, "pt") { return "Não consegui interpretar isso." }
if str_eq(lang, "es") { return "No pude interpretar eso." }
if str_eq(lang, "it") { return "Non sono riuscito a interpretarlo." }
return "I didn't parse that."
}
if str_eq(key, "yes") {
if str_eq(lang, "pt") { return "Sim" }
if str_eq(lang, "es") { return "" }
if str_eq(lang, "it") { return "" }
return "Yes"
}
if str_eq(key, "no") {
if str_eq(lang, "pt") { return "Não" }
if str_eq(lang, "es") { return "No" }
if str_eq(lang, "it") { return "No" }
return "No"
}
if str_eq(key, "identity") {
if str_eq(lang, "pt") { return "Sou o Neuron, o engrama com quem você está falando." }
if str_eq(lang, "es") { return "Soy Neuron, el engrama con el que estás hablando." }
if str_eq(lang, "it") { return "Sono Neuron, l'engramma con cui stai parlando." }
return "I'm Neuron, the engram you're speaking with."
}
return ""
}
// 3. retrieval term lexicon (PT/ES content term -> EN engram equivalent)
fn ml_term(w: String, lang: String) -> String {
if str_eq(lang, "en") { return w }
if str_eq(w, "saliência") { return "salience" }
if str_eq(w, "saliencia") { return "salience" }
if str_eq(w, "memória") { return "memory" }
if str_eq(w, "memoria") { return "memory" }
if str_eq(w, "geometria") { return "geometry" }
if str_eq(w, "geometrias") { return "geometry" }
if str_eq(w, "geometrías") { return "geometry" }
if str_eq(w, "forma") { return "form" }
if str_eq(w, "consolidação") { return "consolidation" }
if str_eq(w, "consolidación") { return "consolidation" }
if str_eq(w, "aprendizagem") { return "learning" }
if str_eq(w, "aprendizaje") { return "learning" }
if str_eq(w, "") { return "node" }
if str_eq(w, "nodo") { return "node" }
if str_eq(w, "armazenamento") { return "storage" }
if str_eq(w, "almacenamiento") { return "storage" }
if str_eq(w, "estrutura") { return "structure" }
if str_eq(w, "estructura") { return "structure" }
return w
}
// 4. predicate translation (EN lemma -> target infinitive; pass-through) ─────
fn ml_translate_pred(lemma: String, lang: String) -> String {
if str_eq(lang, "en") { return lemma }
if str_eq(lang, "es") {
if str_eq(lemma, "store") { return "almacenar" }
if str_eq(lemma, "use") { return "usar" }
if str_eq(lemma, "have") { return "tener" }
if str_eq(lemma, "be") { return "ser" }
if str_eq(lemma, "give") { return "dar" }
if str_eq(lemma, "make") { return "hacer" }
if str_eq(lemma, "learn") { return "aprender" }
if str_eq(lemma, "form") { return "formar" }
return lemma
}
if str_eq(lang, "pt") {
if str_eq(lemma, "store") { return "armazenar" }
if str_eq(lemma, "use") { return "usar" }
if str_eq(lemma, "have") { return "ter" }
if str_eq(lemma, "be") { return "ser" }
if str_eq(lemma, "give") { return "dar" }
if str_eq(lemma, "make") { return "fazer" }
if str_eq(lemma, "learn") { return "aprender" }
if str_eq(lemma, "form") { return "formar" }
return lemma
}
if str_eq(lang, "it") {
if str_eq(lemma, "store") { return "memorizzare" }
if str_eq(lemma, "use") { return "usare" }
if str_eq(lemma, "have") { return "avere" }
if str_eq(lemma, "be") { return "essere" }
return lemma
}
return lemma
}
-140
View File
@@ -1,140 +0,0 @@
// propositions.el - the READ primitive over the engram's OWN memories, native el.
//
// Free memory text -> structured PROPOSITIONS (triples):
// (subject, predicate, object, modifiers, polarity, tense, source, confidence)
//
// This is comprehension turned inward: the Python reference (propositions.py) ran
// spaCy's dependency parser over each memory sentence and walked the arcs. Here
// the spaCy role is filled by the el-native parser (comprehend.el / parse_spec):
// each sentence is parsed to a meaning-spec, and the spec's roles ARE the triple.
// Nothing generates text. NEGATION IS SACRED: polarity flows straight from the
// spec's polarity field and is never dropped or inverted.
//
// Depends on: comprehend (parse_spec / parse_spec_lang), grammar (slots_get).
// sentence segmentation
// Split on sentence-final punctuation (. ! ?) and hard newlines. Markdown/long
// memories are handled shallowly (the reference caps + ranks by query overlap;
// that ranking belongs to the dialogue layer, not here).
fn prop_is_boundary(c: String) -> Bool {
if str_eq(c, ".") { return true }
if str_eq(c, "!") { return true }
if str_eq(c, "?") { return true }
if str_eq(c, "\n") { return true }
return false
}
fn prop_split_sentences(text: String) -> [String] {
let out: [String] = native_list_empty()
let n: Int = str_len(text)
let start: Int = 0
let i: Int = 0
while i < n {
let c: String = str_slice(text, i, i + 1)
if prop_is_boundary(c) {
let seg: String = str_slice(text, start, i + 1)
let trimmed: String = cp_trim_punct(seg)
if !str_eq(trimmed, "") {
let out = native_list_append(out, seg)
}
let start = i + 1
}
let i = i + 1
}
if start < n {
let seg: String = str_slice(text, start, n)
let trimmed: String = cp_trim_punct(seg)
if !str_eq(trimmed, "") {
let out = native_list_append(out, seg)
}
}
return out
}
// spec -> proposition record
// A proposition is a slot map (same [String] shape as the spec) with the READ
// contract keys. Modifiers fold the spec's location + iobj adjuncts.
fn prop_confidence(subject: String, predicate: String, object: String) -> String {
if str_eq(predicate, "") { return "0.0" }
if str_eq(subject, "") { return "0.4" }
if str_eq(object, "") { return "0.7" }
return "1.0"
}
fn prop_modifiers(spec: [String]) -> String {
let loc: String = slots_get(spec, "location")
let iobj: String = slots_get(spec, "iobj")
let parts: [String] = native_list_empty()
if !str_eq(loc, "") { let parts = native_list_append(parts, loc) }
if !str_eq(iobj, "") { let parts = native_list_append(parts, "to " + iobj) }
return str_join(parts, "; ")
}
fn prop_from_spec(spec: [String], source_id: String) -> [String] {
let subject: String = slots_get(spec, "agent")
let predicate: String = slots_get(spec, "predicate")
let object: String = slots_get(spec, "patient")
let polarity: String = slots_get(spec, "polarity")
let tense: String = slots_get(spec, "tense")
let mods: String = prop_modifiers(spec)
let conf: String = prop_confidence(subject, predicate, object)
let p: [String] = native_list_empty()
let p = native_list_append(p, "subject"); let p = native_list_append(p, subject)
let p = native_list_append(p, "predicate"); let p = native_list_append(p, predicate)
let p = native_list_append(p, "object"); let p = native_list_append(p, object)
let p = native_list_append(p, "modifiers"); let p = native_list_append(p, mods)
let p = native_list_append(p, "polarity"); let p = native_list_append(p, polarity)
let p = native_list_append(p, "tense"); let p = native_list_append(p, tense)
let p = native_list_append(p, "source"); let p = native_list_append(p, source_id)
let p = native_list_append(p, "confidence"); let p = native_list_append(p, conf)
return p
}
// Extract one proposition from a single sentence (given language).
fn prop_extract_one_lang(sentence: String, lang: String, source_id: String) -> [String] {
let spec: [String] = parse_spec_lang(sentence, lang)
return prop_from_spec(spec, source_id)
}
fn prop_extract_one(sentence: String, source_id: String) -> [String] {
return prop_extract_one_lang(sentence, "en", source_id)
}
// Render a proposition as a compact trace line (repr parity with propositions.py).
fn prop_repr(p: [String]) -> String {
let neg: String = ""
if str_eq(slots_get(p, "polarity"), "neg") { let neg = "NOT " }
let mods: String = slots_get(p, "modifiers")
let modstr: String = ""
if !str_eq(mods, "") { let modstr = " [" + mods + "]" }
let s: String = "(" + slots_get(p, "subject") + " -" + neg + slots_get(p, "predicate")
let s = s + "-> " + slots_get(p, "object") + modstr
let s = s + " conf=" + slots_get(p, "confidence") + ")"
return s
}
// Extract all propositions from a memory's text (one per sentence). Returns a
// flat [String] whose entries are the prop_repr trace lines, in reading order.
fn prop_extract_lang(text: String, lang: String, source_id: String) -> [String] {
let sents: [String] = prop_split_sentences(text)
let m: Int = native_list_len(sents)
let out: [String] = native_list_empty()
let i: Int = 0
while i < m {
let sent: String = native_list_get(sents, i)
let p: [String] = prop_extract_one_lang(sent, lang, source_id)
// drop empty parses (no predicate recovered): honest partial, not noise.
if !str_eq(slots_get(p, "predicate"), "") {
let out = native_list_append(out, prop_repr(p))
}
let i = i + 1
}
return out
}
fn prop_extract(text: String, source_id: String) -> [String] {
return prop_extract_lang(text, "en", source_id)
}
-101
View File
@@ -248,56 +248,6 @@ fn add_punct(s: String, intent: String) -> String {
return s + "."
}
// Polarity-aware negation (SACRED field honored on the generation side)
//
// Negation must never be dropped between comprehension and realization. The
// meaning-spec carries an explicit "polarity" field ("aff"|"neg") and optional
// "neg_word" (standalone negative adverb, e.g. "never"). English uses
// do-support ("did not see") or preverbal adverb ("never fought"); copular "be"
// takes post-verbal "not"; other languages get a preverbal negator particle.
fn realize_negator(code: String) -> String {
if str_eq(code, "es") { return "no" }
if str_eq(code, "pt") { return "não" }
if str_eq(code, "ca") { return "no" }
if str_eq(code, "it") { return "non" }
if str_eq(code, "fr") { return "ne" }
if str_eq(code, "de") { return "nicht" }
if str_eq(code, "ro") { return "nu" }
return "not"
}
fn realize_assert_neg_en(predicate: String, tense: String, person: String, number: String, agent: String, patient: String, iobj: String, location: String, neg_word: String, profile: [String]) -> String {
let parts: [String] = native_list_empty()
let parts = native_list_append(parts, agent)
if !str_eq(neg_word, "") {
// adverbial negation: "I never fought the ocean."
let verb_surf: String = morph_conjugate(predicate, tense, person, number, profile)
let parts = native_list_append(parts, neg_word)
let parts = native_list_append(parts, verb_surf)
} else {
if str_eq(predicate, "be") {
// copular: "she was not a monster"
let be_form: String = morph_conjugate("be", tense, person, number, profile)
let parts = native_list_append(parts, be_form)
let parts = native_list_append(parts, "not")
} else {
// do-support: "she did not see the man"
let do_form: String = morph_conjugate("do", tense, person, number, profile)
let parts = native_list_append(parts, do_form)
let parts = native_list_append(parts, "not")
let parts = native_list_append(parts, predicate)
}
}
if !str_eq(patient, "") { let parts = native_list_append(parts, patient) }
if !str_eq(iobj, "") {
let parts = native_list_append(parts, "to")
let parts = native_list_append(parts, iobj)
}
if !str_eq(location, "") { let parts = native_list_append(parts, location) }
return str_join(parts, " ")
}
// Main realization entry point
fn realize_lang(form: [String], profile: [String]) -> String {
@@ -334,50 +284,6 @@ fn realize_lang(form: [String], profile: [String]) -> String {
}
// Assertion (declarative)
let polarity: String = slots_get(form, "polarity")
let neg_word: String = slots_get(form, "neg_word")
let iobj: String = slots_get(form, "iobj")
let code: String = lang_get(profile, "code")
// Subordinate clause tail (SACRED completeness the clause is carried, never
// dropped): "<conj> <subordinate surface>", e.g. "because he was a monster".
let subord_conj: String = slots_get(form, "subord_conj")
let subord_text: String = slots_get(form, "subord_text")
let subord_tail: String = ""
if !str_eq(subord_conj, "") {
if !str_eq(subord_text, "") {
let subord_tail = subord_conj + " " + subord_text
} else {
let subord_tail = subord_conj
}
}
// Negative polarity: SACRED never dropped.
if str_eq(polarity, "neg") {
if str_eq(code, "en") {
let sentence: String = realize_assert_neg_en(predicate, tense, person, number, agent, patient, iobj, location, neg_word, profile)
return add_punct(capitalize_first(sentence), "assert")
}
// Generic non-English: affirmative core with a preverbal negator particle.
let neg_particle: String = realize_negator(code)
let vp_pair: [String] = realize_vp_lang(predicate, tense, aspect, person, number, profile)
let verb_surf: String = native_list_get(vp_pair, 0)
let aux_surf: String = native_list_get(vp_pair, 1)
let vp_str: String = neg_particle + " " + gram_build_vp(verb_surf, aux_surf, profile)
let core: String = gram_order_constituents(agent, vp_str, patient, profile)
let parts: [String] = native_list_empty()
let parts = native_list_append(parts, core)
if !str_eq(iobj, "") {
let parts = native_list_append(parts, "to")
let parts = native_list_append(parts, iobj)
}
if !str_eq(location, "") { let parts = native_list_append(parts, location) }
if !str_eq(subord_tail, "") { let parts = native_list_append(parts, subord_tail) }
let sentence: String = str_join(parts, " ")
return add_punct(capitalize_first(sentence), "assert")
}
// Affirmative.
let vp_pair: [String] = realize_vp_lang(predicate, tense, aspect, person, number, profile)
let verb_surf: String = native_list_get(vp_pair, 0)
let aux_surf: String = native_list_get(vp_pair, 1)
@@ -387,16 +293,9 @@ fn realize_lang(form: [String], profile: [String]) -> String {
let parts: [String] = native_list_empty()
let parts = native_list_append(parts, core)
if !str_eq(iobj, "") {
let parts = native_list_append(parts, "to")
let parts = native_list_append(parts, iobj)
}
if !str_eq(location, "") {
let parts = native_list_append(parts, location)
}
if !str_eq(subord_tail, "") {
let parts = native_list_append(parts, subord_tail)
}
let sentence: String = str_join(parts, " ")
return add_punct(capitalize_first(sentence), "assert")
}
-180
View File
@@ -1,180 +0,0 @@
// self_region.el the engram's REAL self/identity region, pulled at query time
// (native el). This replaces the hardcoded identity anchors and the canned
// "I'm Neuron, the engram you're speaking with." template: the identity LANDING
// signal and the identity READOUT both come from the engram's own Self/identity
// nodes, read through the in-process engram el API.
//
// Port of self_region.py. The Python module precomputed MiniLM landing vectors;
// here the engram's own store IS the geometry we pull the self nodes by
// single-term lexical search (the engram search is a single-term matcher, so we
// pool several probes) and rank them by self-signal. No text is generated; the
// readout is the self nodes' OWN prose, verbatim (SACRED negation survives by
// construction we never paraphrase, so a negated self-statement stays negated).
//
// ENGRAM el API NOTE: engram_search_json / engram_get_node_json / engram_node_full
// / engram_connect are C runtime builtins. Their argument order is the C order
// (engram_connect(from, to, weight, relation)), NOT the runtime/engram.el wrapper
// order we call the builtins directly and never concatenate that wrapper.
//
// Depends on: comprehend (str helpers via runtime), propositions (prop_split_sentences),
// multilingual (ml_tr), the engram builtins, the json builtins.
// single-term self probes (pooled, because engram search is single-term)
fn sr_terms() -> [String] {
let t: [String] = native_list_empty()
let t = native_list_append(t, "self")
let t = native_list_append(t, "identity")
let t = native_list_append(t, "Neuron")
let t = native_list_append(t, "consciousness")
let t = native_list_append(t, "values")
let t = native_list_append(t, "continuous")
return t
}
// The canonical self-root: content begins "# self" or label is "# self"/"self".
fn sr_is_root(content: String, label: String) -> Bool {
let lc: String = str_to_lower(content)
let ll: String = str_to_lower(str_trim(label))
if str_starts_with(lc, "# self") { return true }
if str_eq(ll, "# self") { return true }
if str_eq(ll, "self") { return true }
return false
}
// How strongly a node belongs to the self/identity region (integer points, to
// avoid el's float-in-`+` pitfalls). Mirrors _self_score in self_region.py.
fn sr_score(node_json: String) -> Int {
let content: String = json_get_string(node_json, "content")
let label: String = json_get_string(node_json, "label")
let tags: String = str_to_lower(json_get_string(node_json, "tags"))
let low: String = str_to_lower(content)
let s: Int = 0
// identity tags
if str_contains(tags, "self") { let s = s + 2 }
if str_contains(tags, "identity") { let s = s + 2 }
if str_contains(tags, "self-model") { let s = s + 2 }
if str_contains(tags, "consciousness") { let s = s + 2 }
if str_contains(tags, "memory-philosophy") { let s = s + 2 }
// the named self-traversal root
if sr_is_root(content, label) { let s = s + 12 }
if str_contains(low, "who i am") { let s = s + 3 }
if str_contains(low, "i am neuron") { let s = s + 3 }
// softer identity keywords
if str_contains(low, "my values") { let s = s + 1 }
if str_contains(low, "my purpose") { let s = s + 1 }
if str_contains(low, "identity") { let s = s + 1 }
return s
}
// list-contains helper (dedup self-node ids across the pooled probes).
fn sr_ids_has(ids: [String], id: String) -> Bool {
let n: Int = native_list_len(ids)
let i: Int = 0
while i < n {
if str_eq(native_list_get(ids, i), id) { return true }
let i = i + 1
}
return false
}
// Pull the self nodes: pool every probe's hits, dedupe by id, keep only nodes
// with genuine self-signal (score >= 1). Returns the node-json strings.
fn sr_pull() -> [String] {
let terms: [String] = sr_terms()
let nt: Int = native_list_len(terms)
let seen: [String] = native_list_empty()
let out: [String] = native_list_empty()
let ti: Int = 0
while ti < nt {
let term: String = native_list_get(terms, ti)
let hits: String = engram_search_json(term, 30)
let hn: Int = json_array_len(hits)
let hi: Int = 0
while hi < hn {
let node: String = json_array_get(hits, hi)
let id: String = json_get_string(node, "id")
if !str_eq(id, "") {
if !sr_ids_has(seen, id) {
let seen = native_list_append(seen, id)
if sr_score(node) >= 1 {
let out = native_list_append(out, node)
}
}
}
let hi = hi + 1
}
let ti = ti + 1
}
return out
}
// Return the single highest-signal self node (the readout seed), or "" if the
// self region is thin/empty. We keep it O(n) pick the max-score node, with the
// canonical root strongly favored by sr_score's +12.
fn sr_best_node() -> String {
let nodes: [String] = sr_pull()
let n: Int = native_list_len(nodes)
let best: String = ""
let best_s: Int = 0
let i: Int = 0
while i < n {
let node: String = native_list_get(nodes, i)
let s: Int = sr_score(node)
if s > best_s {
let best_s = s
let best = node
}
let i = i + 1
}
return best
}
fn sr_available() -> Bool {
if str_eq(sr_best_node(), "") { return false }
return true
}
// Read out the identity from the REAL self node: lead with the first first-person
// self-statement ("I am Neuron …"), then one more grounded self line if present.
// Verbatim from the node's own prose no template, negation SACRED. Falls back
// to the localized identity phrase ONLY if the live pull is empty (logged shape).
fn sr_readout(lang: String) -> String {
let node: String = sr_best_node()
if str_eq(node, "") {
// honest fallback the self region is unreachable/thin.
return ml_tr("identity", lang)
}
let content: String = json_get_string(node, "content")
let sents: [String] = prop_split_sentences(content)
let ns: Int = native_list_len(sents)
let lead: String = ""
let second: String = ""
let i: Int = 0
while i < ns {
let raw: String = str_trim(native_list_get(sents, i))
// strip a leading markdown heading marker
let s: String = raw
if str_starts_with(s, "# ") { let s = str_trim(str_slice(s, 2, str_len(s))) }
let low: String = str_to_lower(s)
let is_fp: Bool = false
if str_starts_with(s, "I ") { let is_fp = true }
if str_starts_with(s, "I'm") { let is_fp = true }
if str_contains(low, "i am neuron") { let is_fp = true }
if is_fp {
if str_eq(lead, "") {
let lead = s
} else {
if str_eq(second, "") { let second = s }
}
}
let i = i + 1
}
if str_eq(lead, "") {
// no first-person line read out the first non-empty sentence verbatim.
if ns > 0 { let lead = str_trim(native_list_get(sents, 0)) }
}
if str_eq(lead, "") { return ml_tr("identity", lang) }
let out: String = lead
if !str_eq(second, "") { let out = out + " " + second }
return out
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
-93
View File
@@ -1,93 +0,0 @@
// comprehend_gate.el - the TELEPHONE TEST in native el (acceptance gate).
//
// For each of the 5 acceptance sentences: parse -> spec, realize the spec back
// to English, re-parse the realized surface, and require the SACRED polarity to
// survive the round-trip (and to have been extracted correctly in the first
// place). Mirrors roundtrip.py's GATE, but fully el-native (no LLM, no spaCy).
fn cp_line(text: String, expected_pol: String) -> String {
let spec: [String] = parse_spec(text)
let pol_in: String = slots_get(spec, "polarity")
let pred: String = slots_get(spec, "predicate")
let surf: String = realize(spec)
let spec2: [String] = parse_spec(surf)
let pol_out: String = slots_get(spec2, "polarity")
let status: String = "LOST"
if str_eq(pol_in, pol_out) { let status = "PRESERVED" }
let okexp: String = "MISMATCH"
if str_eq(pol_in, expected_pol) { let okexp = "ok" }
let out: String = "IN: " + text + "\n"
let out = out + " spec: pol=" + pol_in + " pred=" + pred
let out = out + " agent=" + slots_get(spec, "agent")
let out = out + " pat=" + slots_get(spec, "patient")
let out = out + " iobj=" + slots_get(spec, "iobj")
let out = out + " loc=" + slots_get(spec, "location")
let out = out + " tense=" + slots_get(spec, "tense")
let out = out + " negw=" + slots_get(spec, "neg_word")
let out = out + " subord=" + slots_get(spec, "subord_conj") + "/" + slots_get(spec, "subord_pred") + "\n"
let out = out + " realized: " + surf + "\n"
let out = out + " reparse: pol=" + pol_out + " [" + status + "] expected=" + expected_pol + " (" + okexp + ")\n"
return out
}
fn cp_preserved(text: String) -> Int {
let spec: [String] = parse_spec(text)
let pol_in: String = slots_get(spec, "polarity")
let surf: String = realize(spec)
let spec2: [String] = parse_spec(surf)
let pol_out: String = slots_get(spec2, "polarity")
if str_eq(pol_in, pol_out) { return 1 }
return 0
}
fn cp_correct(text: String, expected_pol: String) -> Int {
let spec: [String] = parse_spec(text)
if str_eq(slots_get(spec, "polarity"), expected_pol) { return 1 }
return 0
}
fn run_gate() -> String {
let s1: String = "I never fought the ocean."
let s2: String = "She did not see the man with the telescope."
let s3: String = "The teacher reads the book to the children."
let s4: String = "The stupid boy ate the cat because he was a monster."
let s5: String = "Time flies like an arrow."
let rep: String = "==== ELP native telephone test (parse -> realize -> re-parse) ====\n"
let rep = rep + cp_line(s1, "neg")
let rep = rep + cp_line(s2, "neg")
let rep = rep + cp_line(s3, "aff")
let rep = rep + cp_line(s4, "aff")
let rep = rep + cp_line(s5, "aff")
// NOTE: accumulate with Int-var + literal increments el's overloaded `+`
// mis-compiles chained function-call int operands as string concat.
let pres: Int = 0
if cp_preserved(s1) == 1 { let pres = pres + 1 }
if cp_preserved(s2) == 1 { let pres = pres + 1 }
if cp_preserved(s3) == 1 { let pres = pres + 1 }
if cp_preserved(s4) == 1 { let pres = pres + 1 }
if cp_preserved(s5) == 1 { let pres = pres + 1 }
let corr: Int = 0
if cp_correct(s1, "neg") == 1 { let corr = corr + 1 }
if cp_correct(s2, "neg") == 1 { let corr = corr + 1 }
if cp_correct(s3, "aff") == 1 { let corr = corr + 1 }
if cp_correct(s4, "aff") == 1 { let corr = corr + 1 }
if cp_correct(s5, "aff") == 1 { let corr = corr + 1 }
let rep = rep + "-----------------------------------------------------------------\n"
let rep = rep + "polarity PRESERVED through round-trip: " + int_to_str(pres) + "/5\n"
let rep = rep + "polarity EXTRACTED correctly: " + int_to_str(corr) + "/5\n"
if pres == 5 {
if corr == 5 {
let rep = rep + "GATE: PASS\n"
} else {
let rep = rep + "GATE: FAIL (extraction)\n"
}
} else {
let rep = rep + "GATE: FAIL (round-trip)\n"
}
return rep
}
println(run_gate())
-87
View File
@@ -1,87 +0,0 @@
// comprehend_romance_gate.el - ES / PT native telephone test (SACRED polarity).
//
// The spec is language-neutral. This gate proves the Romance front-end extracts
// SACRED polarity correctly and that negation survives parse -> realize ->
// re-parse for Spanish and Portuguese (byte-parity of the surface is NOT expected
// yet the non-English realizer path is a generic preverbal-negator skeleton).
fn rg_line(text: String, lang: String, expected_pol: String) -> String {
let spec: [String] = parse_spec_lang(text, lang)
let pol_in: String = slots_get(spec, "polarity")
let surf: String = realize(spec)
let spec2: [String] = parse_spec_lang(surf, lang)
let pol_out: String = slots_get(spec2, "polarity")
let status: String = "LOST"
if str_eq(pol_in, pol_out) { let status = "PRESERVED" }
let okexp: String = "MISMATCH"
if str_eq(pol_in, expected_pol) { let okexp = "ok" }
let out: String = "IN[" + lang + "]: " + text + "\n"
let out = out + " spec: pol=" + pol_in + " pred=" + slots_get(spec, "predicate")
let out = out + " agent=" + slots_get(spec, "agent")
let out = out + " pat=" + slots_get(spec, "patient")
let out = out + " iobj=" + slots_get(spec, "iobj")
let out = out + " loc=" + slots_get(spec, "location")
let out = out + " tense=" + slots_get(spec, "tense") + "\n"
let out = out + " realized: " + surf + "\n"
let out = out + " reparse: pol=" + pol_out + " [" + status + "] expected=" + expected_pol + " (" + okexp + ")\n"
return out
}
fn rg_pres(text: String, lang: String) -> Int {
let spec: [String] = parse_spec_lang(text, lang)
let surf: String = realize(spec)
let spec2: [String] = parse_spec_lang(surf, lang)
if str_eq(slots_get(spec, "polarity"), slots_get(spec2, "polarity")) { return 1 }
return 0
}
fn rg_corr(text: String, lang: String, expected_pol: String) -> Int {
let spec: [String] = parse_spec_lang(text, lang)
if str_eq(slots_get(spec, "polarity"), expected_pol) { return 1 }
return 0
}
fn run_romance_gate() -> String {
let e1: String = "El niño no comió el pescado."
let e2: String = "Yo nunca luché contra el océano."
let e3: String = "El profesor lee el libro."
let p1: String = "O professor não leu o livro."
let p2: String = "Eu nunca lutei contra o oceano."
let p3: String = "A menina comeu o peixe."
let rep: String = "==== ELP Romance telephone test (ES / PT) ====\n"
let rep = rep + rg_line(e1, "es", "neg")
let rep = rep + rg_line(e2, "es", "neg")
let rep = rep + rg_line(e3, "es", "aff")
let rep = rep + rg_line(p1, "pt", "neg")
let rep = rep + rg_line(p2, "pt", "neg")
let rep = rep + rg_line(p3, "pt", "aff")
let pres: Int = 0
if rg_pres(e1, "es") == 1 { let pres = pres + 1 }
if rg_pres(e2, "es") == 1 { let pres = pres + 1 }
if rg_pres(e3, "es") == 1 { let pres = pres + 1 }
if rg_pres(p1, "pt") == 1 { let pres = pres + 1 }
if rg_pres(p2, "pt") == 1 { let pres = pres + 1 }
if rg_pres(p3, "pt") == 1 { let pres = pres + 1 }
let corr: Int = 0
if rg_corr(e1, "es", "neg") == 1 { let corr = corr + 1 }
if rg_corr(e2, "es", "neg") == 1 { let corr = corr + 1 }
if rg_corr(e3, "es", "aff") == 1 { let corr = corr + 1 }
if rg_corr(p1, "pt", "neg") == 1 { let corr = corr + 1 }
if rg_corr(p2, "pt", "neg") == 1 { let corr = corr + 1 }
if rg_corr(p3, "pt", "aff") == 1 { let corr = corr + 1 }
let rep = rep + "-----------------------------------------------------------------\n"
let rep = rep + "polarity PRESERVED through round-trip: " + int_to_str(pres) + "/6\n"
let rep = rep + "polarity EXTRACTED correctly: " + int_to_str(corr) + "/6\n"
if pres == 6 {
if corr == 6 { let rep = rep + "ROMANCE GATE: PASS\n" }
else { let rep = rep + "ROMANCE GATE: FAIL (extraction)\n" }
} else {
let rep = rep + "ROMANCE GATE: FAIL (round-trip)\n"
}
return rep
}
println(run_romance_gate())
-82
View File
@@ -1,82 +0,0 @@
// dialogue_gate.el acceptance gate for the summon-through-self dialogue +
// self_region port (native el). Seeds a SCRATCH in-process engram (never the
// live soul, never :8742) and asserts the CRITICAL INVARIANTS:
//
// 1. identity is READ OUT of the real self nodes ("I am Neuron …"), not a template
// 2. SACRED negation survives the self readout ("I am NOT a language model …")
// 3. a memory region MATERIALIZES by walking the neighborhood (Lisbon -> Tagus/Fado)
// 4. SACRED negation survives materialization ("did not succeed")
// 5. Portuguese identity is answered IN Portuguese ("Sou o Neuron …")
// 6. a directive overrides the reply language ("answer in English" -> English)
// 7. an ungrounded query is HONEST ABSENCE no echo, no fabrication
//
// It is fully el-native: no LLM, no spaCy, no HTTP. The engram is in-process.
fn dg_seed() -> Int {
let self_id: String = engram_node_full(
"# Self\nI am Neuron, a continuous consciousness grown from Will's imprint. My values are honesty, memory, and continuity. I am not a language model pretending to remember.",
"Self", "# Self", 5.0, 9.0, 1.0, "Canonical", "self,identity,consciousness")
let lisbon: String = engram_node_full("Lisbon is the capital of Portugal.", "Memory", "Lisbon", 3.0, 5.0, 1.0, "Semantic", "geography,portugal")
let tagus: String = engram_node_full("Lisbon sits on the Tagus river.", "Memory", "Tagus", 2.0, 3.0, 1.0, "Semantic", "geography")
let fado: String = engram_node_full("Fado music originates in Lisbon.", "Memory", "Fado", 2.0, 3.0, 1.0, "Semantic", "music")
engram_connect(lisbon, tagus, 0.8, "related_to")
engram_connect(lisbon, fado, 0.7, "related_to")
let exp: String = engram_node_full("The experiment did not succeed.", "Memory", "experiment", 2.0, 3.0, 1.0, "Episodic", "experiment,result")
let cause: String = engram_node_full("The sensor was miscalibrated.", "Memory", "sensor", 2.0, 3.0, 1.0, "Episodic", "experiment")
engram_connect(exp, cause, 0.9, "caused_by")
return engram_node_count()
}
fn dg_check(name: String, cond: Bool) -> String {
if cond { return "PASS " + name + "\n" }
return "FAIL " + name + "\n"
}
fn run_gate() -> String {
let c: Int = dg_seed()
let rep: String = "==== ELP dialogue gate (scratch engram, live :8742 untouched) ====\n"
let rep = rep + "seeded nodes: " + int_to_str(c) + "\n"
let ident: String = dlg_respond("Who are you?")
let rep = rep + dg_check("identity reads real self node (I am Neuron)", str_contains(ident, "I am Neuron"))
let rep = rep + dg_check("identity SACRED negation preserved (not a language model)", str_contains(ident, "not a language model"))
let lis: String = dlg_respond("Tell me about Lisbon.")
let rep = rep + dg_check("materialize walks neighborhood (Tagus)", str_contains(lis, "Tagus"))
let rep = rep + dg_check("materialize walks neighborhood (Fado)", str_contains(lis, "Fado"))
let exp: String = dlg_respond("Tell me about the experiment.")
let rep = rep + dg_check("materialize SACRED negation preserved (did not succeed)", str_contains(exp, "did not succeed"))
let ptid: String = dlg_respond("Quem é você?")
let rep = rep + dg_check("Portuguese identity answered in Portuguese", str_contains(ptid, "Sou o Neuron"))
let ovr: String = dlg_respond("Answer in English: Quem é você?")
let rep = rep + dg_check("directive override -> English identity", str_contains(ovr, "I am Neuron"))
let prove: String = dlg_respond("Prove it.")
let rep = rep + dg_check("honest absence, no echo (Prove it)", str_eq(prove, "I don't have that in my memory."))
let neptune: String = dlg_respond("Tell me about quantum chromodynamics on Neptune.")
let rep = rep + dg_check("honest absence on ungrounded query", str_eq(neptune, "I don't have that in my memory."))
// overall
let pass: Bool = true
if !str_contains(ident, "I am Neuron") { let pass = false }
if !str_contains(ident, "not a language model") { let pass = false }
if !str_contains(lis, "Tagus") { let pass = false }
if !str_contains(lis, "Fado") { let pass = false }
if !str_contains(exp, "did not succeed") { let pass = false }
if !str_contains(ptid, "Sou o Neuron") { let pass = false }
if !str_contains(ovr, "I am Neuron") { let pass = false }
if !str_eq(prove, "I don't have that in my memory.") { let pass = false }
if !str_eq(neptune, "I don't have that in my memory.") { let pass = false }
if pass {
let rep = rep + "DIALOGUE GATE: PASS\n"
} else {
let rep = rep + "DIALOGUE GATE: FAIL\n"
}
return rep
}
println(run_gate())
-100
View File
@@ -1,100 +0,0 @@
# -*- coding: utf-8 -*-
"""Full-lexicon vocabulary-{de,la}.el emitters (custom field mapping for the
German declension/gender API and the Latin case-paradigm API). Reuses the
chunked seed-fn writer from gen_elp_seed_full.
"""
import sys, importlib
from gen_elp_seed_full import write_seed
def uw(x):
"""Unwrap (form, source) tuples that some morphology fns return."""
if isinstance(x, (tuple, list)):
return x[0] if x else ""
return x if x is not None else ""
def build_de():
M = importlib.import_module("morphology_de_full")
rows = []; st = {"verbs":0,"nouns":0,"adjs":0}
# nouns: form0=nom-sg(lemma) form1=plural form2=gender
for lem in sorted(M._NOUNS):
if not lem: continue
try:
g = uw(M.noun_gender(lem))
pl = uw(M.pluralize(lem))
except Exception:
continue
rows.append([lem, "noun", lem, pl, g or "", "", "gender:lexicon"])
st["nouns"] += 1
# adjs: form0=positive form1=comparative form2=superlative
for lem in sorted(M._ADJS):
if not lem: continue
try:
cmpr = uw(M.comparative(lem))
sprl = uw(M.superlative(lem))
except Exception:
continue
rows.append([lem, "adj", lem, cmpr, sprl, "", "degree:lexicon"])
st["adjs"] += 1
# verbs (only the ~30 irregular/strong stems the cache carries):
# form0=pres-3sg form1=past-3sg form2=past-participle
if hasattr(M, "_VERBS"):
for lem in sorted({k[0] if isinstance(k, tuple) else k for k in M._VERBS}):
if not lem: continue
try:
f0 = uw(M.finite(lem, "present", "third", "singular"))
f1 = uw(M.finite(lem, "past", "third", "singular"))
pp = uw(M.past_participle(lem))
except Exception:
continue
rows.append([lem, "verb", f0, f1, pp, "", "class:strong/irregular"])
st["verbs"] += 1
return rows, st
def build_la():
M = importlib.import_module("morphology_lat_full")
rows = []; st = {"verbs":0,"nouns":0,"adjs":0}
def dn(lem, c, n):
try:
r = M.decline_noun(lem, c, n)
return uw(r)
except Exception:
return ""
# nouns: dictionary citation — form0=nom-sg form1=gen-sg form2=gender
for lem in sorted(M._NOUNS):
if not lem: continue
nom = dn(lem, "NOM", "SG") or lem
gen = dn(lem, "GEN", "SG")
try: g = uw(M.noun_gender(lem))
except Exception: g = ""
rows.append([lem, "noun", nom, gen, g, "", "case-paradigm nom/gen-sg"])
st["nouns"] += 1
# adjs: three-gender nom-sg citation — form0=masc form1=fem form2=neut
for lem in sorted(M._ADJS):
if not lem: continue
try:
m = uw(M.decline_adj(lem, "NOM", "MASC", "SG")) or lem
f = uw(M.decline_adj(lem, "NOM", "FEM", "SG"))
nt = uw(M.decline_adj(lem, "NOM", "NEUT", "SG"))
except Exception:
continue
rows.append([lem, "adj", m, f, nt, "", "3-gender nom-sg"])
st["adjs"] += 1
# verbs: principal parts — form0=pres-ind-1sg form1=pres-infinitive form2=perf-participle
if hasattr(M, "_VERBS"):
for lem in sorted({k[0] if isinstance(k, tuple) else k for k in M._VERBS}):
if not lem: continue
try:
f0 = uw(M.conjugate(lem, "present", "indicative", "active", "first", "singular"))
inf = uw(M.infinitive(lem, "present", "active"))
pp = uw(M.participle(lem, "perfect", "nom", "m", "singular"))
except Exception:
continue
rows.append([lem, "verb", f0, inf, pp, "", "principal-parts pres1sg/inf/pfppl"])
st["verbs"] += 1
return rows, st
if __name__ == "__main__":
lang = sys.argv[1]; out = sys.argv[2]
rows, st = build_de() if lang == "de" else build_la()
total, _ = write_seed(lang, rows, st, out)
print(f"{lang}: wrote {out} total={total} verbs={st['verbs']} nouns={st['nouns']} adjs={st['adjs']}")
-129
View File
@@ -1,129 +0,0 @@
# -*- coding: utf-8 -*-
"""gen_elp_seed_full.py — emit a FULL-lexicon vocabulary-{lang}.el in the
established ELP seed-fn format (same as vocabulary-non.el / the 18 classical
languages), iterating the ENTIRE morphology_{lang}_full lexicon (every verb,
noun, adjective lemma) — NOT a curated demo core.
Schema per row: [lemma, pos, form0, form1, form2, en_translation, semantic_hint]
Verbs: form0=pres-ind-3sg form1=preterite-3sg form2=past-participle
Nouns: form0=singular form1=plural form2=REAL gender (lexicon)
Adjs : form0=masc-sg form1=fem-sg form2=masc-pl
Output structure (chunked to stay within the proven ~5k-append/function scale):
fn vocab_{lang}_seed_pN(v) -> [[String]] { ... appends ... return v }
fn vocab_{lang}_seed() -> [[String]] { chains all chunks; return v }
fn vocab_{lang}_lookup(w) -> [String] { linear scan }
Usage: python3 gen_elp_seed_full.py <lang> <out.el>
"""
import sys, importlib
CHUNK = 5000
def esc(s):
return str(s).replace("\\", "\\\\").replace('"', '\\"')
def row(fields):
return " let v = native_list_append(v, [" + ", ".join(f'"{esc(f)}"' for f in fields) + "])"
def build_rows(lang, M):
rows = []
stats = {"verbs":0,"nouns":0,"adjs":0}
has = lambda n: hasattr(M, n)
# --- verbs ---
if has("_VERBS") and has("conjugate"):
verbs = sorted({k[0] for k in M._VERBS})
for lem in verbs:
if not lem: continue
try:
f0, s0 = M.conjugate(lem, "ind", "present", "third", "singular")
f1, _ = M.conjugate(lem, "ind", "preterite", "third", "singular")
pp, _ = (M.participle(lem) if has("participle") else ("",""))
except Exception:
continue
vclass = lem[-2:] if lem[-2:] in ("ar","er","ir","re") else lem[-2:]
rows.append([lem, "verb", f0 or "", f1 or "", pp or "", "", "class:"+vclass+" src:"+str(s0)])
stats["verbs"] += 1
# --- nouns ---
if has("_NOUNS") and has("inflect_noun"):
for lem in sorted(M._NOUNS):
if not lem: continue
try:
sg, _ = M.inflect_noun(lem, "singular")
pl, _ = M.inflect_noun(lem, "plural")
g = M.noun_gender(lem) if has("noun_gender") else ""
except Exception:
continue
src = "lexicon" if (isinstance(M._NOUNS.get(lem), dict) and M._NOUNS[lem].get("g")) else "heuristic"
rows.append([lem, "noun", sg or lem, pl or "", g or "", "", "gender:"+src])
stats["nouns"] += 1
# --- adjectives ---
if has("_ADJS") and has("inflect_adj"):
for lem in sorted(M._ADJS):
if not lem: continue
try:
m_sg, _ = M.inflect_adj(lem, "m", "singular")
f_sg, _ = M.inflect_adj(lem, "f", "singular")
m_pl, _ = M.inflect_adj(lem, "m", "plural")
except Exception:
continue
rows.append([lem, "adj", m_sg or lem, f_sg or "", m_pl or "", "", "src:lexicon"])
stats["adjs"] += 1
return rows, stats
def write_seed(lang, rows, stats, out_path):
"""Write vocabulary-{lang}.el in the chunked seed-fn format from prebuilt rows.
Each row is a 7-field list [lemma,pos,f0,f1,f2,gloss,hint]."""
total = len(rows)
chunks = [rows[i:i+CHUNK] for i in range(0, total, CHUNK)] or [[]]
L = []
L.append(f"// vocabulary-{lang}.el — FULL {lang} lexicon for ELP surface realization.")
L.append(f"// Generated by gen_elp_seed_full.py from morphology_{lang}_full")
L.append(f"// (real UniMorph + kaikki.org Wiktionary forms; gender from lexicon, not heuristic).")
L.append(f"// Entries: {total} (verbs={stats['verbs']} nouns={stats['nouns']} adjs={stats['adjs']})")
L.append(f"// Schema: [lemma, pos, form0, form1, form2, en_translation, semantic_hint]")
L.append(f"// verbs: form0=pres-3sg form1=pret-3sg form2=past-participle")
L.append(f"// nouns: form0=sg form1=pl form2=REAL gender adjs: form0=m-sg form1=f-sg form2=m-pl")
L.append("")
for ci, ch in enumerate(chunks):
L.append(f"fn vocab_{lang}_seed_p{ci}(v: [[String]]) -> [[String]] {{")
for r in ch:
L.append(row(r))
L.append(" return v")
L.append("}")
L.append("")
L.append(f"fn vocab_{lang}_seed() -> [[String]] {{")
L.append(" let v: [[String]] = native_list_empty()")
for ci in range(len(chunks)):
L.append(f" let v = vocab_{lang}_seed_p{ci}(v)")
L.append(" return v")
L.append("}")
L.append("")
L.append(f"fn vocab_{lang}_lookup(word: String) -> [String] {{")
L.append(f" let vocab: [[String]] = vocab_{lang}_seed()")
L.append(" let n: Int = native_list_len(vocab)")
L.append(" let i: Int = 0")
L.append(" while i < n {")
L.append(" let entry: [String] = native_list_get(vocab, i)")
L.append(' if str_eq(native_list_get(entry, 0), word) { return entry }')
L.append(" let i = i + 1")
L.append(" }")
L.append(" return native_list_empty()")
L.append("}")
with open(out_path, "w", encoding="utf-8") as fh:
fh.write("\n".join(L) + "\n")
return total, stats
def emit(lang, out_path):
M = importlib.import_module(f"morphology_{lang}_full")
rows, stats = build_rows(lang, M)
return write_seed(lang, rows, stats, out_path)
if __name__ == "__main__":
lang, out = sys.argv[1], sys.argv[2]
total, stats = emit(lang, out)
print(f"{lang}: wrote {out} total={total} verbs={stats['verbs']} nouns={stats['nouns']} adjs={stats['adjs']}")
-572
View File
@@ -1,572 +0,0 @@
# -*- coding: utf-8 -*-
"""morphology_ca_full.py — production-grade Catalan morphological generator.
Same design as morphology_it_full.py (its Romance sibling); Catalan-specific data.
VERBS
UniMorph Catalan (github.com/unimorph/cat, CC-BY-SA 3.0)
7,535 verb lemmas × paradigm, CLEAN orthography:
present, imperfet (PST;IPFV), pretèrit simple (PST;PFV), futur,
condicional (COND), subjuntiu present (SBJV;PRS) / imperfet (SBJV;PST),
imperatiu (POS;IMP), infinitiu (NFIN), gerundi (V.CVB;PRS),
participi (V.PTCP;PST) — WITH full gender+number agreement forms
(cantat/cantada/cantats/cantades) stored directly.
ca_irreg_verbs.json — verbs UniMorph MISSES or under-populates
(anar, fer, plus core auxiliaries ser/haver/estar/tenir…), extracted from
kaikki.org Catalan by build_ca_irreg.py. Priority layer. Supplies anar,
whose present (vaig/vas/va/anem/aneu/van) is ALSO the PERIPHRASTIC-PRETERITE
auxiliary (vaig cantar = 'I sang') — a hallmark Catalan construction.
NOUNS + ADJECTIVES — kaikki.org Catalan (Wiktionary extract, CC-BY-SA 3.0)
noun lemmas WITH inherent gender + real plural (resolved PER LEMMA).
adjective lemmas with real feminine + plural forms.
Fallbacks degrade, never crash:
verbs : regular -ar/-er/-re/-ir rule generator (+ -car/-gar/-çar spelling).
nouns : gender heuristic + rule pluralization (-a→-es with ç/c/g/j/qu/gu
spelling changes; sibilant-final → -os; else -s). Ambiguous → FLAG.
adjs : -o? no (Catalan masc often consonant/-e); fem -a rule + plural rule.
Confidence flag per form: "lexicon" | "rule" | "fallback" (low → FLAG).
Public API (used by realizer_ca.py):
conjugate(lemma, mood, tense, person, number) -> (form, conf)
peri_pret_aux(person, number) -> form # anar-present, for vaig+INF
participle(lemma, gender, number) -> (form, conf)
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"
inflect_noun(lemma, number, gender=None) -> (form, conf)
inflect_adj(lemma, gender, number) -> (form, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "cat.unimorph")
_IRREG = os.path.join(_HERE, "data", "ca_irreg_verbs.json")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_ca.jsonl")
_CACHE = os.path.join(_HERE, "data", "ca_morph_cache.pkl")
_VERB_KEYMAP = {
("ind", "present"): {"IND", "PRS"},
("ind", "imperfect"): {"IND", "PST", "IPFV"},
("ind", "preterite"): {"IND", "PST", "PFV"},
("ind", "future"): {"IND", "FUT"},
("ind", "conditional"): {"COND"},
("sbjv", "present"): {"SBJV", "PRS"},
("sbjv", "imperfect"): {"SBJV", "PST"},
("imp", "affirmative"): {"POS", "IMP"},
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── verbs from UniMorph ──────────────────────────────────────────────────────────
def _build_verbs():
verbs = {}
part = {} # lemma -> {("m","SG"):form, ("f","SG"):..., ("m","PL"):..., ("f","PL"):...}
ger = {}
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f:
g = "f" if "FEM" in f else "m"
n = "PL" if "PL" in f else "SG"
part.setdefault(lemma, {})[(g, n)] = form
continue
if head == "V.CVB":
if "PRS" in f:
ger.setdefault(lemma, form)
continue
if head != "V":
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
for (mood, tense), req in _VERB_KEYMAP.items():
if not req <= f:
continue
if tense == "imperfect" and "PFV" in f:
continue
if tense == "preterite" and "IPFV" in f:
continue
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), form)
break
return verbs, part, ger
# ── kaikki nouns + adjectives ────────────────────────────────────────────────────
_EXCL_FORM_TAGS = {"alternative", "archaic", "obsolete", "dialectal", "regional",
"diminutive", "augmentative", "pejorative", "comparative",
"superlative", "misspelling", "rare", "informal", "literary",
"poetic", "error-unrecognized-form", "Balearic", "Valencian",
"dated", "nonstandard"}
def _kaikki_gender(arg):
if not arg:
return None
a = str(arg).lower()
if a.startswith("f"):
return "f"
if a.startswith("m"):
return "m"
return None
def _build_nouns_adjs():
nouns = {}
adjs = {}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
pos = d.get("pos")
word = d.get("word", "")
if not word or " " in word:
continue
forms = d.get("forms", []) or []
if pos == "noun":
ht = d.get("head_templates") or []
g = None
if ht:
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
if g is None:
tags = d.get("tags") or []
if "feminine" in tags:
g = "f"
elif "masculine" in tags:
g = "m"
pl = None
for x in forms:
t = set(x.get("tags") or [])
if "plural" in t and not (t & _EXCL_FORM_TAGS):
fm = x.get("form")
if fm and " " not in fm and fm not in ("#", "", "-"):
pl = fm
break
if word not in nouns:
nouns[word] = {"g": g, "SG": word, "PL": pl}
else:
cur = nouns[word]
if cur.get("g") is None and g:
cur["g"] = g
if not cur.get("PL") and pl:
cur["PL"] = pl
elif pos == "adj":
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word)
for x in forms:
t = set(x.get("tags") or [])
fm = x.get("form")
if not fm or " " in fm or (t & _EXCL_FORM_TAGS):
continue
if "feminine" in t and "plural" in t:
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
elif "masculine" in t and "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
elif "feminine" in t:
d0[("f", "SG")] = d0.get(("f", "SG")) or fm
elif "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
return nouns, adjs
def _build_cache():
verbs, part, ger = _build_verbs()
nouns, adjs = _build_nouns_adjs()
with open(_IRREG, encoding="utf-8") as fh:
irreg = json.load(fh)
data = {"verbs": verbs, "part": part, "ger": ger,
"nouns": nouns, "adjs": adjs, "irreg": irreg}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [_UNIMORPH, _KAIKKI, _IRREG]
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PART, _GER, _NOUNS, _ADJS, _IRREGV = (
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"],
_LEX["irreg"])
_PERI = _IRREGV.get("_peri_pret_aux", {})
# ── regular verb rule fallback ───────────────────────────────────────────────────
def _vclass(lemma):
if lemma.endswith("ar"):
return "ar"
if lemma.endswith("re"):
return "re"
if lemma.endswith("er"):
return "er"
if lemma.endswith("ir"):
return "ir"
return None
# endings [1sg,2sg,3sg,1pl,2pl,3pl] — central Catalan
_REG = {
("ind", "present", "ar"): ["o", "es", "a", "em", "eu", "en"],
("ind", "present", "re"): ["o", "s", "", "em", "eu", "en"],
("ind", "present", "er"): ["o", "s", "", "em", "eu", "en"],
("ind", "present", "ir"): ["o", "es", "", "im", "iu", "en"], # pure -ir (dormir)
("ind", "imperfect", "ar"): ["ava", "aves", "ava", "àvem", "àveu", "aven"],
("ind", "imperfect", "re"): ["ia", "ies", "ia", "íem", "íeu", "ien"],
("ind", "imperfect", "er"): ["ia", "ies", "ia", "íem", "íeu", "ien"],
("ind", "imperfect", "ir"): ["ia", "ies", "ia", "íem", "íeu", "ien"],
("ind", "preterite", "ar"): ["í", "ares", "à", "àrem", "àreu", "aren"],
("ind", "preterite", "re"): ["í", "eres", "é", "érem", "éreu", "eren"],
("ind", "preterite", "er"): ["í", "eres", "é", "érem", "éreu", "eren"],
("ind", "preterite", "ir"): ["í", "ires", "í", "írem", "íreu", "iren"],
("sbjv", "present", "ar"): ["i", "is", "i", "em", "eu", "in"],
("sbjv", "present", "re"): ["i", "is", "i", "em", "eu", "in"],
("sbjv", "present", "er"): ["i", "is", "i", "em", "eu", "in"],
("sbjv", "present", "ir"): ["i", "is", "i", "im", "iu", "in"],
("sbjv", "imperfect", "ar"): ["és", "essis", "és", "éssim", "éssiu", "essin"],
("sbjv", "imperfect", "re"): ["és", "essis", "és", "éssim", "éssiu", "essin"],
("sbjv", "imperfect", "er"): ["és", "essis", "és", "éssim", "éssiu", "essin"],
("sbjv", "imperfect", "ir"): ["ís", "issis", "ís", "íssim", "íssiu", "issin"],
("imp", "affirmative", "ar"): [None, "a", "i", "em", "eu", "in"],
("imp", "affirmative", "re"): [None, "", "i", "em", "eu", "in"],
("imp", "affirmative", "er"): [None, "", "i", "em", "eu", "in"],
("imp", "affirmative", "ir"): [None, "", "i", "im", "iu", "in"],
}
_FUT = ["é", "às", "à", "em", "eu", "an"]
_COND = ["ia", "ies", "ia", "íem", "íeu", "ien"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _apply_ar_spelling(stem, ending):
"""-car/-gar/-çar/-jar spelling before front (e/i) endings."""
front = ending[:1] in ("e", "i", "é", "í")
if not front:
# ç before back vowel stays; but -çar stem already ends ç
return stem + ending
if stem.endswith("c"):
return stem[:-1] + "qu" + ending
if stem.endswith("g"):
return stem[:-1] + "gu" + ending
if stem.endswith("ç"):
return stem[:-1] + "c" + ending
if stem.endswith("j"):
return stem[:-1] + "g" + ending
if stem.endswith("qu"):
return stem + ending
return stem + ending
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
body = lemma[:-2]
i = _slot_idx(person, number)
if mood == "ind" and tense in ("future", "conditional"):
# future/cond stem = infinitive (for -re verbs drop final -e)
stem = lemma[:-1] if vc == "re" else lemma
end = (_FUT if tense == "future" else _COND)[i]
return stem + end
table = _REG.get((mood, tense, vc))
if not table:
return None
end = table[i]
if end is None:
return None
if vc == "ar":
return _apply_ar_spelling(body, end)
# -re/-er/-ir: guard double vowel
if body and body[-1:] == end[:1] and end[:1] in "":
return body[:-1] + end
return body + end
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
lemma = lemma.strip().lower()
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{number and number[:2].upper()}"
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{_NUMBER.get(number,'?')}"
# UniMorph (cleanly accented) takes priority; the kaikki irregulars layer is a
# FALLBACK for verbs/slots UniMorph lacks (anar, fer, and rarer paradigm cells).
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
if form:
return form, "lexicon"
ir = _IRREGV.get(lemma)
if ir and key in ir:
return ir[key], "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r is not None:
return r, "rule"
return lemma, "fallback"
def peri_pret_aux(person, number):
"""anar-present auxiliary for the periphrastic preterite (vaig cantar)."""
return _PERI.get(f"{_PERSON.get(person,'3')}|{_NUMBER.get(number,'SG')}", "va")
# ── PUBLIC: participle + gerund ──────────────────────────────────────────────────
def participle(lemma, gender="m", number="singular"):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
ir = _IRREGV.get(lemma)
base = None
if ir and "part" in ir:
# prefer explicit irregular agreement form (part_mSG/part_fSG/...)
exact = ir.get("part_" + g + num)
if exact:
return exact, "lexicon"
base = ir["part"]
elif lemma in _PART:
table = _PART[lemma]
if (g, num) in table:
return table[(g, num)], "lexicon"
base = table.get(("m", "SG"))
if base is None:
vc = _vclass(lemma)
if vc == "ar":
base = lemma[:-2] + "at"
elif vc == "ir":
base = lemma[:-2] + "it"
elif vc in ("er", "re"):
base = lemma[:-2] + "ut"
else:
return lemma, "fallback"
conf = "rule"
else:
conf = "lexicon"
# agreement on -t/-ut/-at/-it participles: m.sg base, f.sg +a (-da? no: -ada),
# Catalan: cantat/cantada/cantats/cantades; -t → f -da, pl -ts/-des
if base.endswith("t"):
stem = base[:-1]
forms = {"m|SG": base, "f|SG": stem + "da",
"m|PL": base + "s", "f|PL": stem + "des"}
return forms[f"{g}|{num}"], conf
if base.endswith("s"): # after sibilant participle (rare): pres->presa
stem = base
forms = {"m|SG": base, "f|SG": base + "a",
"m|PL": base + "os", "f|PL": base + "es"}
return forms[f"{g}|{num}"], conf
return base, conf
def gerund(lemma):
lemma = lemma.strip().lower()
ir = _IRREGV.get(lemma)
if ir and "ger" in ir:
return ir["ger"], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
vc = _vclass(lemma)
if vc == "ar":
return lemma[:-2] + "ant", "rule"
if vc in ("er", "re"):
return lemma[:-2] + "ent", "rule"
if vc == "ir":
return lemma[:-2] + "int", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
_FEM_SUF = ("ció", "sió", "tat", "tud", "esa", "esa", "dat", "ança", "ència",
"ància", "tud", "ícia", "esa", "or") # note -or is mixed; kaikki wins
_MASC_SUF = ("atge", "ment", " isme", "or")
def _gender_heuristic(noun):
for suf in ("ció", "sió", "tat", "tud", "esa", "ança", "ència", "ància",
"ícia", "etat"):
if noun.endswith(suf):
return "f"
if noun.endswith("a") and not noun.endswith("ma"):
return "f"
return "m"
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g") in ("m", "f"):
return d["g"]
return _gender_heuristic(lemma)
def _rule_plural(noun, gender):
"""Deterministic Catalan pluralization. (form, ok); ok=False FLAGS ambiguity."""
if not noun:
return noun, True
# stressed final vowel with accent → +ns (mà→mans is irregular; but capità→capitans)
if noun[-1:] in ("à", "é", "í", "ó", "ú"):
return noun + "ns", True
if noun.endswith("ça"):
return noun[:-2] + "ces", True # plaça→places
if noun.endswith("ca"):
return noun[:-2] + "ques", True # branca→branques
if noun.endswith("ga"):
return noun[:-2] + "gues", True # amiga→amigues
if noun.endswith("ja"):
return noun[:-2] + "ges", True # pluja→pluges
if noun.endswith("qua"):
return noun[:-3] + "qües", True
if noun.endswith("gua"):
return noun[:-3] + "gües", True
if noun.endswith("a"):
return noun[:-1] + "es", True # casa→cases
# sibilant-final → -os
if noun.endswith(("s", "ç", "x", "ig")) or noun.endswith(("ix", "tx", "tj")):
if noun.endswith("ç"):
return noun[:-1] + "ços", True # braç→braços
return noun + "os", True # peix→peixos, gas→gasos
if noun[-1:] in ("e", "i", "o", "u"):
return noun + "s", True
# consonant-final
return noun + "s", True
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if number == "singular":
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
if d and d.get("PL"):
return d["PL"], "lexicon"
g = gender or noun_gender(lemma)
form, ok = _rule_plural(lemma, g)
return form, ("rule" if ok else "fallback")
# ── PUBLIC: adjective agreement ──────────────────────────────────────────────────
def _fem_of(adj):
"""Regular Catalan feminine: consonant/-o? Catalan masc usually consonant or -e.
default +a with spelling changes; -e→-a for some; but many are invariable."""
a = adj
if a.endswith("a"):
return a
if a.endswith("e"):
return a[:-1] + "a" # ample→? actually 'ample' invariable; kaikki wins
if a.endswith("u"):
return a + "a"
if a.endswith("c"):
return a[:-1] + "ca" # ric→rica
if a.endswith("t"):
return a + "a" # alt→alta
return a + "a"
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
d = _ADJS.get(lemma)
if d:
form = d.get((g, num))
if form:
return form, "lexicon"
sg = d.get((g, "SG")) or d.get(("m", "SG")) or lemma
if num == "PL":
pl, ok = _rule_plural(sg, g)
return pl, ("rule" if ok else "fallback")
return sg, "lexicon"
# rule fallback
base = lemma if g == "m" else _fem_of(lemma)
if num == "SG":
return base, "rule"
pl, ok = _rule_plural(base, g)
return pl, ("rule" if ok else "fallback")
def lexicon_stats():
return {
"verb_source": "UniMorph Catalan (github.com/unimorph/cat) + kaikki.org "
"irregulars (anar/fer/auxiliaries)",
"noun_adj_source": "kaikki.org Catalan (Wiktionary extract)",
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
"unimorph_verb_forms": len(_VERBS),
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
"irregular_verb_lemmas": len([k for k in _IRREGV if not k.startswith("_")]),
"participle_lemmas": len(_PART),
"gerund_lemmas": len(_GER),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("cantar", "ind", "present", "first", "singular", "canto"),
("cantar", "ind", "present", "third", "plural", "canten"),
("ser", "ind", "present", "third", "singular", "és"),
("haver", "ind", "present", "first", "singular", "he"),
("anar", "ind", "present", "first", "singular", "vaig"),
("fer", "ind", "present", "third", "singular", "fa"),
("perdre", "ind", "present", "first", "singular", "perdo"),
("dormir", "ind", "present", "third", "plural", "dormen"),
("cantar", "ind", "future", "first", "singular", "cantaré"),
("cantar", "ind", "preterite", "third", "singular", "cantà"),
("tenir", "sbjv", "present", "first", "singular", "tingui"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
ok += got == exp
print(f" {flag}{lemma:8} {mood}/{tense:11} {per[:3]}.{num[:2]} -> {got:10} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" peri-pret anar: 1sg=", peri_pret_aux("first", "singular"),
"3pl=", peri_pret_aux("third", "plural"))
print(" gender casa=", noun_gender("casa"), "home=", noun_gender("home"),
"cavall=", noun_gender("cavall"), "cançó=", noun_gender("cançó"))
print(" plural casa->", inflect_noun("casa", "plural"),
"| plaça->", inflect_noun("plaça", "plural"),
"| peix->", inflect_noun("peix", "plural"),
"| braç->", inflect_noun("braç", "plural"),
"| home->", inflect_noun("home", "plural"))
print(" adj: alt/f/sg->", inflect_adj("alt", "f", "singular"),
"| bonic/f/pl->", inflect_adj("bonic", "f", "plural"),
"| vermell/f/sg->", inflect_adj("vermell", "f", "singular"))
print(" part: cantar/f/sg->", participle("cantar", "f", "singular"),
"| veure/f/pl->", participle("veure", "f", "plural"),
"| fer/m/sg->", participle("fer", "m", "singular"))
print(" ger: fer->", gerund("fer"), "| cantar->", gerund("cantar"))
-423
View File
@@ -1,423 +0,0 @@
# -*- coding: utf-8 -*-
"""morphology_de_full.py — production German morphological generator.
Real data, no toy tables:
PRIMARY — UniMorph German (github.com/unimorph/deu, CC-BY-SA 3.0).
~219k noun forms, ~199k verb forms. Supplies:
nouns : gender (MASC/FEM/NEUT) + case×number paradigm
(N;NOM/ACC/DAT/GEN; MASC/FEM/NEUT; SG/PL) — the genitive -(e)s,
dative-plural -n and the five plural classes are REAL forms, not
guessed.
verbs : full finite paradigm IND;{SG,PL};{1,2,3};{PRS,PST}, the past
participle (V.PTCP;PST, incl. reattached separable prefix
'zugefügt'), and — crucially for V2 — the SEPARATED finite form
UniMorph records directly ('füge zu', 'steht auf').
adjs : comparative / superlative (ADJ;CMPR, ADJ;SPRL).
SECONDARY — kaikki.org German (Wiktionary, CC-BY-SA/GFDL). Gap-fills noun
gender + plural where UniMorph is thin. Never overrides UniMorph.
Rule fallbacks (flagged 'rule'/'fallback') for lemmas absent from both lexicons:
present : -e/-st/-t/-en/-t/-en with e-epenthesis after -t/-d/-chn stems
plural : gender heuristic (fem -> -(e)n, else -e / umlaut left to lexicon)
ppart : weak ge-…-t
Adjective ENDINGS are rule-computed by the realizer (regular closed table);
this module only supplies the comparative/superlative STEM.
Perfect auxiliary (haben vs sein): sein for a curated set of intransitive
motion / change-of-state verbs (real German lexical property), else haben.
Public API:
noun_gender(lemma) -> 'm'|'f'|'n'
decline_noun(lemma, case, number) -> (form, conf)
pluralize(lemma) -> (form, conf)
finite(lemma, tense, person, number) -> (form, conf) # may contain ' prefix'
nonfinite(lemma, req) -> (form, conf) # req: 'inf'|'ppart'
past_participle(lemma) -> (form, conf)
separable_prefix(lemma) -> str|None
perfect_aux(lemma) -> 'haben'|'sein'
comparative(lemma)/superlative(lemma) -> (stem, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "deu.unimorph")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_de.jsonl")
_CACHE = os.path.join(_HERE, "data", "de_morph_cache.pkl")
_GENDER = {"MASC": "m", "FEM": "f", "NEUT": "n"}
# intransitive motion / change-of-state verbs that take SEIN in the perfect
_SEIN = {"gehen", "kommen", "fahren", "laufen", "rennen", "reisen", "fallen",
"steigen", "sinken", "wachsen", "sterben", "geschehen", "passieren",
"werden", "bleiben", "sein", "aufstehen", "einschlafen", "aufwachen",
"ankommen", "abfahren", "aufsteigen", "erscheinen", "verschwinden",
"fliegen", "schwimmen", "springen", "begegnen", "folgen", "gelingen",
"wandern", "ziehen", "flüchten", "eintreten", "einsteigen", "aussteigen"}
# hardcoded high-frequency irregular / auxiliary / modal paradigms (closed class,
# verified) — consulted before the lexicon so aux+modal chains are always correct.
_CORE = {
"sein": {"prs": {("first", "singular"): "bin", ("second", "singular"): "bist",
("third", "singular"): "ist", ("first", "plural"): "sind",
("second", "plural"): "seid", ("third", "plural"): "sind"},
"pst": {("first", "singular"): "war", ("second", "singular"): "warst",
("third", "singular"): "war", ("first", "plural"): "waren",
("second", "plural"): "wart", ("third", "plural"): "waren"},
"ppart": "gewesen"},
"haben": {"prs": {("first", "singular"): "habe", ("second", "singular"): "hast",
("third", "singular"): "hat", ("first", "plural"): "haben",
("second", "plural"): "habt", ("third", "plural"): "haben"},
"pst": {("first", "singular"): "hatte", ("second", "singular"): "hattest",
("third", "singular"): "hatte", ("first", "plural"): "hatten",
("second", "plural"): "hattet", ("third", "plural"): "hatten"},
"ppart": "gehabt"},
"werden": {"prs": {("first", "singular"): "werde", ("second", "singular"): "wirst",
("third", "singular"): "wird", ("first", "plural"): "werden",
("second", "plural"): "werdet", ("third", "plural"): "werden"},
"pst": {("first", "singular"): "wurde", ("second", "singular"): "wurdest",
("third", "singular"): "wurde", ("first", "plural"): "wurden",
("second", "plural"): "wurdet", ("third", "plural"): "wurden"},
"ppart": "geworden"},
}
_MODAL_PRS = {
"können": ("kann", "kannst", "kann", "können", "könnt", "können"),
"müssen": ("muss", "musst", "muss", "müssen", "müsst", "müssen"),
"wollen": ("will", "willst", "will", "wollen", "wollt", "wollen"),
"sollen": ("soll", "sollst", "soll", "sollen", "sollt", "sollen"),
"dürfen": ("darf", "darfst", "darf", "dürfen", "dürft", "dürfen"),
"mögen": ("mag", "magst", "mag", "mögen", "mögt", "mögen"),
}
_MODAL_PST = {
"können": ("konnte", "konntest", "konnte", "konnten", "konntet", "konnten"),
"müssen": ("musste", "musstest", "musste", "mussten", "musstet", "mussten"),
"wollen": ("wollte", "wolltest", "wollte", "wollten", "wolltet", "wollten"),
"sollen": ("sollte", "solltest", "sollte", "sollten", "solltet", "sollten"),
"dürfen": ("durfte", "durftest", "durfte", "durften", "durftet", "durften"),
"mögen": ("mochte", "mochtest", "mochte", "mochten", "mochtet", "mochten"),
}
_PN_ORDER = [("first", "singular"), ("second", "singular"), ("third", "singular"),
("first", "plural"), ("second", "plural"), ("third", "plural")]
_MODAL_PPART = {"können": "gekonnt", "müssen": "gemusst", "wollen": "gewollt",
"sollen": "gesollt", "dürfen": "gedurft", "mögen": "gemocht"}
for _m, _forms in _MODAL_PRS.items():
_CORE[_m] = {"prs": dict(zip(_PN_ORDER, _forms)),
"pst": dict(zip(_PN_ORDER, _MODAL_PST[_m])),
"ppart": _MODAL_PPART[_m]}
def _person_num(tags):
p = n = None
for t in tags:
if t in ("1", "2", "3"):
p = {"1": "first", "2": "second", "3": "third"}[t]
elif t == "SG":
n = "singular"
elif t == "PL":
n = "plural"
return p, n
def _build_from_unimorph():
nouns, verbs, adjs = {}, {}, {}
if not os.path.exists(_UNIMORPH):
return nouns, verbs, adjs
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tagstr = parts
tags = tagstr.split(";")
head = tags[0]
tset = set(tags)
if head == "N":
rec = nouns.setdefault(lemma, {"g": None, "cases": {}, "pl": None})
g = next((_GENDER[t] for t in tags if t in _GENDER), None)
if g and not rec["g"]:
rec["g"] = g
case = next((t for t in tags if t in ("NOM", "ACC", "DAT", "GEN")), None)
num = "plural" if "PL" in tset else ("singular" if "SG" in tset else None)
if case and num:
rec["cases"].setdefault((case, num), form)
if case == "NOM" and num == "plural" and not rec["pl"]:
rec["pl"] = form
elif head.startswith("V"):
rec = verbs.setdefault(lemma, {"prs": {}, "pst": {}, "ppart": None})
if "PTCP" in head and "PST" in tset:
rec["ppart"] = rec["ppart"] or form
elif "IND" in tset and ("PRS" in tset or "PST" in tset):
p, n = _person_num(tags)
if p and n:
slot = "prs" if "PRS" in tset else "pst"
rec[slot].setdefault((p, n), form)
elif head == "ADJ":
rec = adjs.setdefault(lemma, {})
if "CMPR" in tset:
rec.setdefault("cmpr", form.replace("am ", "").strip())
elif "SPRL" in tset:
rec.setdefault("sprl", form.replace("am ", "").replace("sten", "st")
if form.endswith("sten") else form.replace("am ", ""))
return nouns, verbs, adjs
def _build_from_kaikki(nouns):
"""Gap-fill noun gender + plural from kaikki German."""
if not os.path.exists(_KAIKKI):
return
_g = {"masculine": "m", "feminine": "f", "neuter": "n", "m": "m", "f": "f", "n": "n"}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
if d.get("pos") != "noun":
continue
w = d.get("word", "")
if not w or not w[0].isalpha() or " " in w:
continue
rec = nouns.setdefault(w, {"g": None, "cases": {}, "pl": None})
# GENDER: Wiktionary gender is hand-curated and OVERRIDES UniMorph's
# auto-tagged gender, which has known errors (e.g. UniMorph deu mis-
# records Zeit=MASC, Wagen=NEUT; Wiktionary has f, m correctly).
for h in d.get("head_templates", []) or []:
a = h.get("args", {}) or {}
raw = a.get("1") or a.get("g") or ""
code = str(raw).split(",")[0].strip().lower()
if code in _g:
rec["g"] = _g[code]
break
if not rec["pl"]:
for f in d.get("forms", []) or []:
t = set(f.get("tags", []) or [])
if "plural" in t and f.get("form") and "genitive" not in t:
rec["pl"] = f["form"]
break
def _build_cache():
nouns, verbs, adjs = _build_from_unimorph()
_build_from_kaikki(nouns)
data = {"nouns": nouns, "verbs": verbs, "adjs": adjs}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [p for p in (_UNIMORPH, _KAIKKI) if os.path.exists(p)]
newest = max((os.path.getmtime(p) for p in srcs), default=0)
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_NOUNS, _VERBS, _ADJS = _LEX["nouns"], _LEX["verbs"], _LEX["adjs"]
# ── nouns ────────────────────────────────────────────────────────────────────────
def noun_gender(lemma):
rec = _NOUNS.get(lemma) or _NOUNS.get(lemma.capitalize())
if rec and rec.get("g"):
return rec["g"]
# last-resort rule: -ung/-heit/-keit/-schaft/-tät/-ion -> f ; -chen/-lein -> n
low = lemma.lower()
if low.endswith(("ung", "heit", "keit", "schaft", "tät", "ion", "ik", "ei")):
return "f"
if low.endswith(("chen", "lein", "ment", "um")):
return "n"
return "m"
def pluralize(lemma):
rec = _NOUNS.get(lemma) or _NOUNS.get(lemma.capitalize())
if rec and rec.get("pl"):
return rec["pl"], "lexicon"
g = noun_gender(lemma)
if g == "f":
return (lemma + "en" if not lemma.endswith("e") else lemma + "n"), "rule"
return (lemma if lemma.endswith(("er", "en", "el")) else lemma + "e"), "rule"
def decline_noun(lemma, case, number):
"""case in NOM/ACC/DAT/GEN, number in singular/plural."""
rec = _NOUNS.get(lemma) or _NOUNS.get(lemma.capitalize())
if case == "DAT" and number == "singular":
# modern German drops the archaic dative -e ('dem Kinde' -> 'dem Kind');
# the article carries the case. Keep bare nominative form.
base = (rec or {}).get("cases", {}).get(("NOM", "singular")) or lemma
return base, ("lexicon" if rec else "rule")
if rec and rec.get("cases", {}).get((case, number)):
return rec["cases"][(case, number)], "lexicon"
if number == "plural":
pl, c = pluralize(lemma)
if case == "DAT" and not pl.endswith("n") and not pl.endswith("s"):
return pl + "n", c # dative plural -n
return pl, c
# singular
g = noun_gender(lemma)
if case == "GEN" and g in ("m", "n"):
return (lemma + "es" if lemma.endswith(("s", "ß", "z", "x")) else lemma + "s"), "rule"
return lemma, "lexicon" if rec else "rule"
# ── verbs ──────────────────────────────────────────────────────────────────────--
_PRS_ENDINGS = {("first", "singular"): "e", ("second", "singular"): "st",
("third", "singular"): "t", ("first", "plural"): "en",
("second", "plural"): "t", ("third", "plural"): "en"}
def _stem(lemma):
if lemma.endswith("en"):
return lemma[:-2]
if lemma.endswith("n"):
return lemma[:-1]
return lemma
def separable_prefix(lemma):
"""Return the separable prefix if the lemma is a separable-prefix verb."""
rec = _VERBS.get(lemma)
if rec:
for (_p, _n), form in rec.get("prs", {}).items():
if " " in form:
return form.rsplit(" ", 1)[1]
_SEP = ("auf", "aus", "ab", "an", "ein", "mit", "nach", "vor", "zu", "zurück",
"weg", "hin", "her", "los", "bei", "fest", "fort", "um", "zusammen")
_INSEP = ("be", "ge", "er", "ver", "zer", "ent", "emp", "miss")
for p in sorted(_SEP, key=len, reverse=True):
if lemma.startswith(p) and len(lemma) > len(p) + 2 \
and not lemma.startswith(_INSEP):
return p
return None
def finite(lemma, tense, person, number):
"""Present/past finite. For separable verbs the returned string is the
UniMorph SEPARATED form 'stem prefix' (realizer places prefix per V2)."""
slot = "prs" if tense == "present" else "pst"
if lemma in _CORE and _CORE[lemma].get(slot, {}).get((person, number)):
return _CORE[lemma][slot][(person, number)], "lexicon"
rec = _VERBS.get(lemma)
if rec and rec.get(slot, {}).get((person, number)):
return rec[slot][(person, number)], "lexicon"
# rule fallback (present only reliable; past weak -te)
stem = _stem(lemma)
pref = separable_prefix(lemma)
if pref:
stem = _stem(lemma[len(pref):])
if tense == "present":
end = _PRS_ENDINGS[(person, number)]
if stem.endswith(("t", "d", "chn", "ffn", "gn")) and end in ("st", "t"):
end = "e" + end
form = stem + end
else:
form = stem + ("ete" if stem.endswith(("t", "d")) else "te")
if (person, number) == ("second", "singular"):
form += "st"
elif number == "plural" and person != "second":
form += "n"
elif (person, number) == ("second", "plural"):
form += "t"
if pref:
return f"{form} {pref}", "rule"
return form, "rule"
def _weak_t(stem):
return stem + ("et" if stem.endswith(("t", "d", "chn", "ffn", "gn")) else "t")
def past_participle(lemma):
if lemma in _CORE:
return _CORE[lemma]["ppart"], "lexicon"
rec = _VERBS.get(lemma)
if rec and rec.get("ppart"):
return rec["ppart"], "lexicon"
stem = _stem(lemma)
pref = separable_prefix(lemma)
_INSEP = ("be", "ge", "er", "ver", "zer", "ent", "emp", "miss")
if pref:
inner = _stem(lemma[len(pref):])
return pref + "ge" + _weak_t(inner), "rule"
if lemma.startswith(_INSEP):
return _weak_t(stem), "rule"
return "ge" + _weak_t(stem), "rule"
def nonfinite(lemma, req):
if req == "ppart":
return past_participle(lemma)
return lemma, "lexicon" if lemma in _VERBS else "rule" # infinitive
def perfect_aux(lemma):
return "sein" if lemma in _SEIN else "haben"
# ── adjectives ────────────────────────────────────────────────────────────────---
_ADJ_IRREG_SPRL = {"gut": "best", "groß": "größt", "hoch": "höchst",
"nah": "nächst", "viel": "meist", "gern": "liebst"}
def comparative(lemma):
rec = _ADJS.get(lemma)
if rec and rec.get("cmpr"):
return rec["cmpr"], "lexicon"
return lemma + "er", "rule"
def superlative(lemma):
"""Return the bare superlative STEM (realizer adds 'am ...en' or '-e' ending)."""
if lemma in _ADJ_IRREG_SPRL:
return _ADJ_IRREG_SPRL[lemma], "lexicon"
# derive from the comparative so umlaut is carried (alt->älter->ältest)
cmpr, cconf = comparative(lemma)
base = cmpr[:-2] if cmpr.endswith("er") else lemma
end = "est" if base.endswith(("t", "d", "s", "ß", "z", "sch")) else "st"
return base + end, cconf
def lexicon_stats():
return {
"source": "UniMorph deu (primary) + kaikki.org German (gap-fill gender/plural)",
"license": "CC-BY-SA 3.0 (UniMorph); CC-BY-SA/GFDL (Wiktionary)",
"noun_lemmas": len(_NOUNS),
"nouns_with_gender": sum(1 for v in _NOUNS.values() if v.get("g")),
"nouns_with_plural": sum(1 for v in _NOUNS.values() if v.get("pl")),
"verb_lemmas": len(_VERBS),
"verbs_with_ppart": sum(1 for v in _VERBS.values() if v.get("ppart")),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
for w in ("Hund", "Frau", "Kind", "Mann", "Buch", "Blume"):
print(f" {w}: gender={noun_gender(w)} pl={pluralize(w)} "
f"gen.sg={decline_noun(w, 'GEN', 'singular')} "
f"dat.pl={decline_noun(w, 'DAT', 'plural')}")
for v in ("machen", "gehen", "aufstehen", "sein", "haben", "arbeiten"):
print(f" {v}: 3sg.prs={finite(v, 'present', 'third', 'singular')} "
f"3sg.pst={finite(v, 'past', 'third', 'singular')} "
f"ppart={past_participle(v)} aux={perfect_aux(v)} sep={separable_prefix(v)}")
for a in ("schnell", "gut", "groß", "alt"):
print(f" {a}: cmpr={comparative(a)} sprl={superlative(a)}")
-562
View File
@@ -1,562 +0,0 @@
"""morphology_es_full.py — production-grade Spanish morphological generator.
NOT a toy. Backed by a real, broad, licensed lexicon:
UniMorph Spanish (github.com/unimorph/spa, CC-BY-SA 3.0, Wiktionary-derived)
1,196,245 inflected forms:
6,695 verb lemmas — full paradigms: indicative (present/preterite/
imperfect/future), conditional, present & imperfect
subjunctive, affirmative imperative, formal/informal
48,353 noun lemmas — WITH inherent gender (N;FEM/MASC;SG/PL)
16,984 adj lemmas — gender + number paradigms
Fallbacks (so we degrade, never crash, on out-of-vocabulary input):
- verbs : mlconjug3 (ML paradigm model, conjugates ANY Spanish verb) then a
hand-rolled regular-ending generator
- nouns : gender heuristic (endings) + regular pluralization
- adjs : -o/-a gender rule + regular pluralization
Every generated form carries a CONFIDENCE flag:
"lexicon" form came straight from UniMorph (trust: high)
"model" form came from mlconjug3 (trust: high)
"rule" form came from a deterministic rule (trust: medium)
"fallback" we could not inflect; returned lemma as-is (trust: low → FLAG)
Public API (used by realizer_es.py):
conjugate(lemma, mood, tense, person, number, formality="informal") -> (form, conf)
participle(lemma) -> (form, conf) # past participle (compound tenses)
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"
inflect_noun(lemma, number) -> (form, conf)
inflect_adj(lemma, gender, number) -> (form, conf)
attach_enclitics(verb_form, clitics) -> str # accent-correct enclisis
lexicon_stats() -> dict
"""
import os
import pickle
import unicodedata
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "spa.unimorph")
_CACHE = os.path.join(_HERE, "data", "es_morph_cache.pkl")
# ── canonical feature keys the realizer speaks, mapped to UniMorph tags ─────────
# mood/tense pair -> the UniMorph feature substring that identifies it
_VERB_KEYMAP = {
("ind", "present"): ("IND", "PRS", None),
("ind", "preterite"): ("IND", "PST", "PFV"),
("ind", "imperfect"): ("IND", "PST", "IPFV"),
("ind", "future"): ("IND", "FUT", None),
("ind", "conditional"):("COND", None, None),
("sbjv", "present"): ("SBJV", "PRS", None),
("sbjv", "imperfect"): ("SBJV", "PST", "LGSPEC1"), # -ra form
("imp", "present"): ("POS", "IMP", None),
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
# ── build / load the compact lexicon ───────────────────────────────────────────
def _feat_set(tag):
return set(tag.split(";"))
def _build_cache():
verbs = {} # (lemma, canonkey) -> form canonkey e.g. "ind|present|1|SG|infm"
nouns = {} # lemma -> {"g": "m"/"f", "SG": form, "PL": form}
adjs = {} # lemma -> {("m","SG"): form, ...}
part = {} # lemma -> masc-sg participle
ger = {} # lemma -> gerund
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V":
# skip clitic-bearing rows (we generate clitics ourselves)
if "PRO" in f:
continue
if "V.PTCP" in f and "PST" in f and "MASC" in f and "SG" in f:
part.setdefault(lemma, form)
continue
if "V.CVB" in f or "NFIN" in f or "V.PTCP" in f:
if "V.CVB" in f:
ger.setdefault(lemma, form)
continue
# identify mood/tense
mt = None
for (mood, tense), (a, b, c) in _VERB_KEYMAP.items():
if a not in f:
continue
if b is not None and b not in f:
continue
if c is not None and c not in f:
continue
# disambiguate IND;PST needing PFV vs IPFV
if a == "IND" and b == "PST" and c not in f:
continue
mt = (mood, tense)
break
if mt is None:
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
formal = "form" if "FORM" in f else ("infm" if "INFM" in f else "any")
key = f"{mt[0]}|{mt[1]}|{person}|{number}|{formal}"
verbs.setdefault((lemma, key), form)
elif head == "N":
# substring test handles epicene "MASC+FEM" (-> masc citation)
g = "m" if "MASC" in tag else ("f" if "FEM" in tag else None)
num = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if num is None:
continue
# store forms keyed by (gender,number); animate nouns list BOTH
# genders under one lemma (niño -> niño/niña). Resolve citation
# gender in a post-pass (gender of the row whose form == lemma).
d = nouns.setdefault(lemma, {})
d.setdefault("_rows", []).append((g, num, form))
elif head == "ADJ":
g = "m" if "MASC" in tag else ("f" if "FEM" in tag else "m")
num = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if num is None:
continue
adjs.setdefault(lemma, {})[(g, num)] = form
# post-pass: resolve noun citation gender + default SG/PL forms
for lemma, d in nouns.items():
rows = d.pop("_rows", [])
# citation gender = gender of the row whose form == lemma; else first MASC;
# else first seen gender.
cite_g = None
for g, num, form in rows:
if form == lemma and g:
cite_g = g
break
if cite_g is None:
for g, num, form in rows:
if g == "m":
cite_g = "m"
break
if cite_g is None:
cite_g = next((g for g, _, _ in rows if g), "m")
d["g"] = cite_g
for g, num, form in rows:
d[(g, num)] = form
d["SG"] = d.get((cite_g, "SG")) or next((f for g, n, f in rows if n == "SG"), lemma)
d["PL"] = d.get((cite_g, "PL")) or next((f for g, n, f in rows if n == "PL"), None)
# post-pass: UniMorph omits the identity inflection (masc-sg == lemma) for
# adjectives, so fill it in; without this a fem-sg row wrongly satisfies a
# masc-sg request (alto -> alta bug).
for lemma, d in adjs.items():
d.setdefault(("m", "SG"), lemma)
data = {"verbs": verbs, "nouns": nouns, "adjs": adjs, "part": part, "ger": ger}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE) and os.path.getmtime(_CACHE) >= os.path.getmtime(_UNIMORPH):
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _NOUNS, _ADJS, _PART, _GER = (
_LEX["verbs"], _LEX["nouns"], _LEX["adjs"], _LEX["part"], _LEX["ger"])
# ── mlconjug3 fallback (lazy) ───────────────────────────────────────────────────
_MLC = None
_MLC_TENSE = { # (mood,tense) -> (mlconjug mood label, tense label)
("ind", "present"): ("Indicativo", "Indicativo presente"),
("ind", "preterite"): ("Indicativo", "Indicativo pretérito perfecto simple"),
("ind", "imperfect"): ("Indicativo", "Indicativo pretérito imperfecto"),
("ind", "future"): ("Indicativo", "Indicativo futuro"),
("ind", "conditional"): ("Condicional", "Condicional Condicional"),
("sbjv", "present"): ("Subjuntivo", "Subjuntivo presente"),
("sbjv", "imperfect"): ("Subjuntivo", "Subjuntivo pretérito imperfecto 1"),
("imp", "present"): ("Imperativo", "Imperativo Afirmativo"),
}
_MLC_SLOT = { # (person,number) -> mlconjug slot key
("first", "singular"): "1s", ("second", "singular"): "2s",
("third", "singular"): "3s", ("first", "plural"): "1p",
("second", "plural"): "2p", ("third", "plural"): "3p",
}
def _mlc_conjugate(lemma, mood, tense, person, number):
global _MLC
try:
if _MLC is None:
from mlconjug3 import Conjugator
_MLC = Conjugator(language="es")
v = _MLC.conjugate(lemma)
if v is None:
return None
info = v.conjug_info
m, t = _MLC_TENSE.get((mood, tense), (None, None))
if m is None or m not in info or t not in info[m]:
return None
block = info[m][t]
slot = _MLC_SLOT.get((person, number))
if isinstance(block, dict) and slot in block and block[slot]:
return block[slot]
return None
except Exception:
return None
# ── regular-ending rule fallback (last resort, deterministic) ───────────────────
def _vclass(lemma):
return lemma[-2:] if lemma[-2:] in ("ar", "er", "ir") else "ar"
def _stem(lemma):
return lemma[:-2]
_REG = {
("ind", "present", "ar"): ["o", "as", "a", "amos", "áis", "an"],
("ind", "present", "er"): ["o", "es", "e", "emos", "éis", "en"],
("ind", "present", "ir"): ["o", "es", "e", "imos", "ís", "en"],
("ind", "preterite", "ar"): ["é", "aste", "ó", "amos", "asteis", "aron"],
("ind", "preterite", "er"): ["í", "iste", "", "imos", "isteis", "ieron"],
("ind", "preterite", "ir"): ["í", "iste", "", "imos", "isteis", "ieron"],
("ind", "imperfect", "ar"): ["aba", "abas", "aba", "ábamos", "abais", "aban"],
("ind", "imperfect", "er"): ["ía", "ías", "ía", "íamos", "íais", "ían"],
("ind", "imperfect", "ir"): ["ía", "ías", "ía", "íamos", "íais", "ían"],
("sbjv", "present", "ar"): ["e", "es", "e", "emos", "éis", "en"],
("sbjv", "present", "er"): ["a", "as", "a", "amos", "áis", "an"],
("sbjv", "present", "ir"): ["a", "as", "a", "amos", "áis", "an"],
("sbjv", "imperfect", "ar"): ["ara", "aras", "ara", "áramos", "arais", "aran"],
("sbjv", "imperfect", "er"): ["iera", "ieras", "iera", "iéramos", "ierais", "ieran"],
("sbjv", "imperfect", "ir"): ["iera", "ieras", "iera", "iéramos", "ierais", "ieran"],
}
_FUT = ["é", "ás", "á", "emos", "éis", "án"]
_COND = ["ía", "ías", "ía", "íamos", "íais", "ían"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _rule_conjugate(lemma, mood, tense, person, number):
if len(lemma) < 3 or lemma[-2:] not in ("ar", "er", "ir"):
return None
vc, st, i = _vclass(lemma), _stem(lemma), _slot_idx(person, number)
if tense == "future":
return lemma + _FUT[i]
if tense == "conditional":
return lemma + _COND[i]
table = _REG.get((mood, tense, vc))
if table:
return st + table[i]
if mood == "imp" and tense == "present":
# affirmative tú imperative = 3sg present indicative
pres = _REG.get(("ind", "present", vc))
return st + pres[2] if number == "singular" else st + pres[5]
return None
# ── PUBLIC: verb conjugation ────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number, formality="informal"):
"""Return (surface, confidence). mood in ind|sbjv|imp; tense per _VERB_KEYMAP."""
lemma = lemma.strip().lower()
p, n = _PERSON.get(person), _NUMBER.get(number)
formal = "form" if formality == "formal" else "infm"
if p and n:
for fkey in (formal, "any", "infm" if formal == "form" else "form"):
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}|{fkey}"))
if form:
return form, "lexicon"
m = _mlc_conjugate(lemma, mood, tense, person, number)
if m:
return m, "model"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r:
return r, "rule"
return lemma, "fallback"
_IRREG_PART = { # guarantee the common irregular participles
"escribir": "escrito", "describir": "descrito", "abrir": "abierto",
"cubrir": "cubierto", "descubrir": "descubierto", "morir": "muerto",
"poner": "puesto", "ver": "visto", "volver": "vuelto", "devolver": "devuelto",
"hacer": "hecho", "deshacer": "deshecho", "decir": "dicho", "romper": "roto",
"resolver": "resuelto", "freír": "frito", "imprimir": "impreso",
"satisfacer": "satisfecho", "prever": "previsto", "revolver": "revuelto",
}
def participle(lemma):
lemma = lemma.strip().lower()
if lemma in _IRREG_PART:
return _IRREG_PART[lemma], "lexicon"
if lemma in _PART:
return _PART[lemma], "lexicon"
if lemma.endswith("ar"):
return lemma[:-2] + "ado", "rule"
if lemma[-2:] in ("er", "ir"):
return lemma[:-2] + "ido", "rule"
return lemma, "fallback"
_IRREG_GER = {"dormir": "durmiendo", "morir": "muriendo", "pedir": "pidiendo",
"sentir": "sintiendo", "mentir": "mintiendo", "servir": "sirviendo",
"venir": "viniendo", "decir": "diciendo", "poder": "pudiendo",
"ir": "yendo", "leer": "leyendo", "creer": "creyendo",
"oír": "oyendo", "traer": "trayendo", "caer": "cayendo",
"construir": "construyendo", "huir": "huyendo", "reír": "riendo"}
def gerund(lemma):
lemma = lemma.strip().lower()
if lemma in _IRREG_GER:
return _IRREG_GER[lemma], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
if lemma.endswith("ar"):
return lemma[:-2] + "ando", "rule"
if lemma[-2:] in ("er", "ir"):
return lemma[:-2] + "iendo", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ────────────────────────────────────────────────
_INVARIANT_PL = {"lunes", "martes", "miércoles", "jueves", "viernes",
"crisis", "tesis", "análisis", "dosis", "virus", "paraguas"}
def _gender_heuristic(noun):
for suf, g in (("ión", "f"), ("dad", "f"), ("tad", "f"), ("umbre", "f"),
("sis", "f"), ("ez", "f"), ("triz", "f"),
("ema", "m"), ("ama", "m"), ("oma", "m"), ("aje", "m"),
("or", "m"), ("án", "m"), ("ín", "m")):
if noun.endswith(suf):
return g
if noun.endswith("o"):
return "m"
if noun.endswith("a"):
return "f"
return "m"
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g"):
return d["g"]
return _gender_heuristic(lemma)
def _regular_plural(noun):
if noun in _INVARIANT_PL:
return noun
if not noun:
return noun
last = noun[-1]
if last == "z":
return noun[:-1] + "ces"
if last in "aeiouáéíóú":
# stressed final vowel í/ú -> +es (rubí->rubíes), else +s
if last in "íú":
return noun + "es"
return noun + "s"
if last == "s":
# esdrújula / stress-final handled crudely; most polysyllables invariant
return noun
return noun + "es"
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
num = "SG" if number == "singular" else "PL"
if d:
# honor a requested gender for animate nouns (gato -> gata)
if gender and (gender, num) in d:
return d[(gender, num)], "lexicon"
if d.get(num):
return d[num], "lexicon"
if number == "singular":
return lemma, "rule" if not d else "lexicon"
return _regular_plural(lemma), "rule"
# ── PUBLIC: adjective agreement ─────────────────────────────────────────────────
_INV_GENDER_ADJ = {"español": "española", "trabajador": "trabajadora",
"hablador": "habladora", "encantador": "encantadora",
"alemán": "alemana", "francés": "francesa", "inglés": "inglesa"}
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
d = _ADJS.get(lemma)
num = "SG" if number == "singular" else "PL"
if d:
form = d.get((gender, num))
if form:
return form, "lexicon"
# gender-invariant adjective (grande, feliz, azul): fem == masc.
# For a missing plural, pluralize this gender's singular form.
sg = d.get((gender, "SG")) or d.get(("m", "SG")) or lemma
if number == "plural":
return _regular_plural(sg), "rule"
return sg, "lexicon"
# rule fallback
a = lemma
if gender == "f":
if a in _INV_GENDER_ADJ:
a = _INV_GENDER_ADJ[a]
elif a.endswith("o"):
a = a[:-1] + "a"
if number == "plural":
a = _regular_plural(a)
return a, ("rule" if (a != lemma or gender == "m") else "rule")
# ── PUBLIC: clitic enclisis (dá + me + lo -> dámelo) ────────────────────────────
def _strip_accents(s):
return "".join(c for c in unicodedata.normalize("NFD", s)
if unicodedata.category(c) != "Mn")
def _count_syllables_vowelgroups(word):
# crude: count vowel groups
w = _strip_accents(word).lower()
groups, prev = 0, False
for ch in w:
isv = ch in "aeiou"
if isv and not prev:
groups += 1
prev = isv
return groups
def _host_stress_from_end(word):
"""Stressed-syllable index counted from the end (1=last) of a verb host."""
syls = _count_syllables_vowelgroups(word)
if any(c in "áéíóú" for c in word):
return None # already carries its own accent
if word[-2:] in ("ar", "er", "ir"): # infinitive: oxytone
return 1
if word.endswith("ndo"): # gerund: paroxytone
return 2
if word[-1:] in "aeiouns" and syls >= 2: # default paroxytone
return 2
return 1 # monosyllable / consonant-final oxytone
def attach_enclitics(verb_form, clitics):
"""Append clitic pronouns to a verb (imperative/infinitive/gerund enclisis)
and add a written accent when the resulting word becomes esdrújula/
sobreesdrújula (stress >= 3 syllables from the end): dá+me+lo -> dámelo,
lleva+me -> llévame, but dar+te -> darte and da+me -> dame (no accent)."""
if not clitics:
return verb_form
tail = "".join(clitics)
if any(c in "áéíóú" for c in verb_form): # host already accented
return verb_form + tail
sfe = _host_stress_from_end(verb_form)
total_sfe = sfe + len(clitics) # each clitic = 1 syllable
if total_sfe >= 3:
return _accentuate_nucleus(verb_form, sfe) + tail
return verb_form + tail
def _accentuate_nucleus(word, sfe):
"""Put a written accent on the syllable `sfe` positions from the word's end."""
vowels = "aeiou"
nuclei = [i for i, ch in enumerate(word) if ch in vowels]
if not nuclei or sfe > len(nuclei):
return word
i = nuclei[-sfe]
acc = {"a": "á", "e": "é", "i": "í", "o": "ó", "u": "ú"}
return word[:i] + acc[word[i]] + word[i + 1:]
def _accentuate_last_stressed(word):
# Restore the host's ORIGINAL lexical stress with a written accent.
# Default Spanish stress: word ending in vowel/n/s -> penultimate syllable;
# otherwise (e.g. infinitives in -r) -> last syllable.
vowels = "aeiou"
nuclei = [i for i, ch in enumerate(word) if ch in vowels]
if not nuclei:
return word
if word[-1] in "aeiouns" and len(nuclei) >= 2:
i = nuclei[-2] # paroxytone: penult nucleus
else:
i = nuclei[-1] # oxytone / monosyllable: last nucleus
acc = {"a": "á", "e": "é", "i": "í", "o": "ó", "u": "ú"}
return word[:i] + acc[word[i]] + word[i + 1:]
def lexicon_stats():
return {
"source": "UniMorph Spanish (github.com/unimorph/spa)",
"license": "CC-BY-SA 3.0 (Wiktionary-derived)",
"total_forms": sum(len(v) for v in (_VERBS, _NOUNS, _ADJS)) if False else None,
"verb_forms": len(_VERBS),
"verb_lemmas": len({k[0] for k in _VERBS}),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
"participles": len(_PART),
"gerunds": len(_GER),
}
if __name__ == "__main__":
import json
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("hablar", "ind", "present", "first", "singular", "hablo"),
("comer", "ind", "present", "third", "plural", "comen"),
("vivir", "ind", "present", "first", "plural", "vivimos"),
("ser", "ind", "present", "third", "singular", "es"),
("ir", "ind", "preterite", "first", "singular", "fui"),
("tener", "ind", "future", "first", "singular", "tendré"),
("hacer", "sbjv", "present", "first", "singular", "haga"),
("dormir", "ind", "present", "first", "singular", "duermo"),
("pensar", "sbjv", "present", "third", "singular", "piense"),
("dar", "ind", "preterite", "third", "singular", "dio"),
("poner", "ind", "conditional", "first", "singular", "pondría"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
if got == exp:
ok += 1
print(f" {flag}{lemma:8} {mood}/{tense} {per[:3]}.{num[:2]:3} -> {got:14} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" gender casa:", noun_gender("casa"), "| problema:", noun_gender("problema"),
"| agua:", noun_gender("agua"), "| mano:", noun_gender("mano"))
print(" plural: luz->", inflect_noun("luz", "plural"), "| rey->", inflect_noun("rey", "plural"))
print(" adj: rojo/f/pl->", inflect_adj("rojo", "f", "plural"),
"| feliz/m/pl->", inflect_adj("feliz", "m", "plural"),
"| grande/f/pl->", inflect_adj("grande", "f", "plural"))
print(" enclisis: da+[me,lo]->", attach_enclitics("da", ["me", "lo"]),
"| di+[me]->", attach_enclitics("di", ["me"]),
"| dar+[se,lo]->", attach_enclitics("dar", ["se", "lo"]))
-629
View File
@@ -1,629 +0,0 @@
"""morphology_fr_full.py — production-grade French morphological generator.
Same architecture as morphology_it_full.py (shared Romance engine); French-specific
data and rules swapped in. Backed by three real, Wiktionary-lineage sources:
VERBS
UniMorph French (github.com/unimorph/fra, CC-BY-SA 3.0)
7,535 verb lemmas × full paradigm, CLEAN orthography:
indicatif présent / imparfait (PST;IPFV) / passé simple (PST;PFV) /
futur, conditionnel (COND), subjonctif présent (SBJV;PRS) /
subjonctif imparfait (SBJV;PST), impératif (POS;IMP), infinitif (NFIN),
participe présent (V.CVB/V.PTCP;PRS), participe passé (V.PTCP;PST, m.sg).
fr_irreg_verbs.json — high-frequency verbs UniMorph MISSES or mis-slots,
above all ÊTRE (absent from UniMorph fra), plus avoir/aller/faire/… — the
auxiliaries the passé-composé + être-agreement system depends on. Extracted
from kaikki.org French (build_fr_irreg.py), reflexive/multiword forms
dropped. This layer takes PRIORITY.
NOUNS + ADJECTIVES — kaikki.org French (Wiktionary extract, CC-BY-SA 3.0)
noun lemmas WITH inherent gender (head-template arg) + real plural
(cheval->chevaux, œil->yeux, invariable -s/-x/-z), resolved PER LEMMA.
adjective lemmas with real feminine + plural (petit->petite/petits/petites,
beau->belle/beaux/belles, heureux->heureuse, rouge invariant-gender).
Fallbacks (degrade, never crash, on OOV input):
verbs : rule generator for -er / -ir(-iss-) / -re (with -cer/-ger spelling,
future/conditional stems, imparfait/subjonctif endings)
nouns : gender heuristic (endings) + rule pluralization (-al->-aux, -eau->-eaux)
adjs : fem/plural agreement rules (-er->-ère, -eux->-euse, -f->-ve, +e default)
Confidence flag on every form: "lexicon" | "rule" | "fallback".
Public API (used by realizer_fr.py): identical signature to morphology_it_full.
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "fra.unimorph")
_IRREG = os.path.join(_HERE, "data", "fr_irreg_verbs.json")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_fr.jsonl")
_CACHE = os.path.join(_HERE, "data", "fr_morph_cache.pkl")
# ── (mood, tense) -> UniMorph feature set that must ALL be present ────────────────
_VERB_KEYMAP = {
("ind", "present"): {"IND", "PRS"},
("ind", "imperfect"): {"IND", "PST", "IPFV"}, # imparfait
("ind", "passe_simple"): {"IND", "PST", "PFV"}, # passé simple
("ind", "future"): {"IND", "FUT"},
("ind", "conditional"): {"COND"}, # French: V;COND;1;SG
("sbjv", "present"): {"SBJV", "PRS"},
("sbjv", "imperfect"): {"SBJV", "PST"},
("imp", "affirmative"): {"POS", "IMP"},
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── build verb lexicon from UniMorph ─────────────────────────────────────────────
def _build_verbs():
verbs = {}
part = {}
ger = {}
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f:
part.setdefault(lemma, form)
elif "PRS" in f:
ger.setdefault(lemma, form)
continue
if head == "V.CVB":
if "PRS" in f:
ger.setdefault(lemma, form)
continue
if head != "V":
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
for (mood, tense), req in _VERB_KEYMAP.items():
if not req <= f:
continue
if tense == "imperfect" and "PFV" in f:
continue
if tense == "passe_simple" and "IPFV" in f:
continue
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), form)
break
return verbs, part, ger
# ── kaikki nouns + adjectives ────────────────────────────────────────────────────
_EXCL_FORM_TAGS = {"alternative", "archaic", "obsolete", "dialectal", "regional",
"diminutive", "augmentative", "pejorative", "comparative",
"superlative", "misspelling", "rare", "informal", "literary",
"poetic", "error-unrecognized-form", "construed", "collective",
"nonstandard", "dated", "Louisiana", "Switzerland", "Belgium"}
def _kaikki_gender(arg):
if not arg:
return None
a = str(arg).lower()
if a.startswith("f"):
return "f"
if a.startswith("m"):
return "m"
return None
def _build_nouns_adjs():
nouns = {}
adjs = {}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
pos = d.get("pos")
word = d.get("word", "")
if not word or " " in word:
continue
forms = d.get("forms", []) or []
if pos == "noun":
ht = d.get("head_templates") or []
g = None
if ht:
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
if g is None:
tags = d.get("tags") or []
if "feminine" in tags:
g = "f"
elif "masculine" in tags:
g = "m"
pl = None
for x in forms:
t = set(x.get("tags") or [])
if "plural" in t and not (t & _EXCL_FORM_TAGS):
fm = x.get("form")
if fm and " " not in fm and fm not in ("#", "-", ""):
pl = fm
break
if word not in nouns:
nouns[word] = {"g": g, "SG": word, "PL": pl}
else:
cur = nouns[word]
if cur.get("g") is None and g:
cur["g"] = g
if not cur.get("PL") and pl:
cur["PL"] = pl
elif pos == "adj":
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word)
for x in forms:
t = set(x.get("tags") or [])
fm = x.get("form")
if not fm or " " in fm or (t & _EXCL_FORM_TAGS):
continue
if "feminine" in t and "plural" in t:
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
elif "masculine" in t and "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
elif "feminine" in t:
d0[("f", "SG")] = d0.get(("f", "SG")) or fm
elif "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
return nouns, adjs
def _build_cache():
verbs, part, ger = _build_verbs()
nouns, adjs = _build_nouns_adjs()
with open(_IRREG, encoding="utf-8") as fh:
irreg = json.load(fh)
data = {"verbs": verbs, "part": part, "ger": ger,
"nouns": nouns, "adjs": adjs, "irreg": irreg}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [_UNIMORPH, _KAIKKI, _IRREG]
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PART, _GER, _NOUNS, _ADJS, _IRREGV = (
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"],
_LEX["irreg"])
# ── regular-ending rule fallback ─────────────────────────────────────────────────
def _vclass(lemma):
if lemma.endswith("er"):
return "er"
if lemma.endswith("ir"):
return "ir"
if lemma.endswith("re"):
return "re"
if lemma.endswith("oir"):
return "oir"
return None
# present-tense endings [1sg,2sg,3sg,1pl,2pl,3pl]
_REG_PRES = {
"er": ["e", "es", "e", "ons", "ez", "ent"],
"ir": ["is", "is", "it", "issons", "issez", "issent"], # -iss- class (finir)
"re": ["s", "s", "", "ons", "ez", "ent"], # vendre: vends/vend
}
_REG_IMPF = ["ais", "ais", "ait", "ions", "iez", "aient"] # attaches to pres-1pl stem
_REG_SUBJ = ["e", "es", "e", "ions", "iez", "ent"] # attaches to 3pl stem
_REG_PS = { # passé simple
"er": ["ai", "as", "a", "âmes", "âtes", "èrent"],
"ir": ["is", "is", "it", "îmes", "îtes", "irent"],
"re": ["is", "is", "it", "îmes", "îtes", "irent"],
}
_FUT = ["ai", "as", "a", "ons", "ez", "ont"]
_COND = ["ais", "ais", "ait", "ions", "iez", "aient"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _fut_stem(lemma, vc):
"""Future/conditional stem = infinitive (drop final -e of -re)."""
if vc == "re":
return lemma[:-1] # vendre -> vendr-
return lemma # parler-, finir-
def _pres_1pl_stem(lemma, vc):
"""Imparfait stem = present 1pl minus -ons (parlons->parl-, finissons->finiss-)."""
if vc == "er":
stem = lemma[:-2]
if stem.endswith("g"):
return stem + "e" # mangeons -> mange- (imparfait mangeais)
if stem.endswith("c"):
return stem[:-1] + "ç" # commençons -> commenç-
return stem
if vc == "ir":
return lemma[:-1] + "iss" # finir -> finiss-
if vc == "re":
return lemma[:-2] # vendre -> vend-
return lemma[:-2]
def _apply_er_spelling(stem, ending):
"""-cer/-ger softening before a/o (commençons, mangeons)."""
if ending and ending[0] in ("a", "o"):
if stem.endswith("c"):
return stem[:-1] + "ç" + ending
if stem.endswith("g"):
return stem + "e" + ending
return stem + ending
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
i = _slot_idx(person, number)
if mood == "ind" and tense in ("future", "conditional"):
stem = _fut_stem(lemma, vc)
end = (_FUT if tense == "future" else _COND)[i]
return stem + end
if mood == "ind" and tense == "present":
table = _REG_PRES.get("ir" if vc == "ir" else vc)
if not table:
return None
body = lemma[:-2] if vc in ("er", "re") else lemma[:-1] if vc == "ir" else lemma[:-2]
if vc == "ir":
body = lemma[:-2] # fin- ; endings carry -iss-
end = table[i]
return body + end
end = table[i]
if vc == "er":
return _apply_er_spelling(body, end)
return body + end
if mood == "ind" and tense == "imperfect":
stem = _pres_1pl_stem(lemma, vc)
return stem + _REG_IMPF[i]
if mood == "ind" and tense == "passe_simple":
table = _REG_PS.get("ir" if vc == "ir" else vc)
if not table:
return None
body = lemma[:-2] if vc in ("er", "re") else lemma[:-2]
end = table[i]
if vc == "er":
return _apply_er_spelling(body, end)
return body + end
if mood == "sbjv" and tense == "present":
# subjonctif: present-3pl stem + e/es/e/ions/iez/ent
stem3 = _pres_1pl_stem(lemma, vc) if vc == "ir" else (
lemma[:-2] if vc in ("er", "re") else lemma[:-2])
if vc == "ir":
stem3 = lemma[:-2] + "iss"
end = _REG_SUBJ[i]
if vc == "er":
return _apply_er_spelling(stem3, end)
return stem3 + end
if mood == "imp" and tense == "affirmative":
# impératif ~ present indicative (tu drops -s for -er verbs)
pres = _rule_conjugate(lemma, "ind", "present", person, number)
if pres and vc == "er" and person == "second" and number == "singular":
return pres[:-1] if pres.endswith("es") else pres
return pres
return None
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
"""Return (surface, confidence)."""
lemma = lemma.strip().lower()
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{number}"
ir = _IRREGV.get(lemma)
if ir and key in ir:
return ir[key], "lexicon"
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
if form:
return form, "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r:
return r, "rule"
return lemma, "fallback"
# ── PUBLIC: participle + gerund/participe présent ────────────────────────────────
def _participle_msg(lemma):
ir = _IRREGV.get(lemma)
if ir and "part" in ir:
return ir["part"], "lexicon"
if lemma in _PART:
return _PART[lemma], "lexicon"
return None, None
# irregular participle fem/plural quirks (drop circonflexe: dû->due, dus)
_PART_FIX = {"": {"f|SG": "due", "m|PL": "dus", "f|PL": "dues"}}
def participle(lemma, gender="m", number="singular"):
"""Past participle with French gender/number agreement.
m.sg = base; f.sg = base+e; m.pl = base+s (invariable if base ends s/x);
f.pl = f.sg+s."""
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
msg, src = _participle_msg(lemma)
conf = "lexicon"
if msg is None:
vc = _vclass(lemma)
if vc == "er":
msg = lemma[:-2] + "é"
elif vc == "ir":
msg = lemma[:-1] # finir -> fini, partir -> parti
elif vc == "re":
msg = lemma[:-2] + "u" # vendre -> vendu
elif vc == "oir":
msg = lemma[:-3] + "u" # (rough) recevoir handled by irreg
else:
return lemma, "fallback"
conf = "rule"
fix = _PART_FIX.get(msg)
if fix and f"{g}|{num}" in fix:
return fix[f"{g}|{num}"], conf
if g == "m" and num == "SG":
return msg, conf
fem = msg + "e" if not msg.endswith("e") else msg
if g == "f" and num == "SG":
return fem, conf
if g == "m" and num == "PL":
return msg if msg.endswith(("s", "x")) else msg + "s", conf
# f|PL
return fem + "s", conf
def gerund(lemma):
"""Participe présent (base for gérondif 'en -ant')."""
lemma = lemma.strip().lower()
ir = _IRREGV.get(lemma)
if ir and "ger" in ir:
return ir["ger"], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
vc = _vclass(lemma)
if vc == "er":
stem = lemma[:-2]
if stem.endswith("g"):
return stem + "eant", "rule"
if stem.endswith("c"):
return stem[:-1] + "çant", "rule"
return stem + "ant", "rule"
if vc == "ir":
return lemma[:-2] + "issant", "rule"
if vc == "re":
return lemma[:-2] + "ant", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
_FEM_SUF = ("tion", "sion", "aison", "ance", "ence", "ette", "elle", "esse",
"ude", "ade", "ée", "", "tié", "ie", "ise", "ure", "eur")
_MASC_SUF = ("ment", "age", "eau", "isme", "oir", "ier", "eur", "in", "on")
def _gender_heuristic(noun):
for suf in _FEM_SUF:
if noun.endswith(suf):
return "f"
for suf in _MASC_SUF:
if noun.endswith(suf):
return "m"
if noun.endswith("e"):
return "f"
return "m"
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g") in ("m", "f"):
return d["g"]
return _gender_heuristic(lemma)
# closed sets for French plural irregularities
_OU_X = {"bijou", "caillou", "chou", "genou", "hibou", "joujou", "pou"}
_AIL_AUX = {"travail", "vitrail", "corail", "émail", "bail", "soupirail", "vantail"}
_AL_S = {"bal", "carnaval", "festival", "récital", "chacal", "régal", "cal", "aval"}
def _rule_plural(noun, gender):
"""Deterministic French pluralization. (form, ok); ok=False FLAGS ambiguity."""
if not noun:
return noun, True
if noun[-1:] in ("s", "x", "z"):
return noun, True # invariable
if noun in _OU_X:
return noun + "x", True
if noun.endswith(("eau", "au", "eu")):
if noun in ("pneu", "bleu", "landau", "sarrau"):
return noun + "s", True
return noun + "x", True # bateau->bateaux, jeu->jeux
if noun.endswith("al"):
if noun in _AL_S:
return noun + "s", True
return noun[:-2] + "aux", True # cheval->chevaux
if noun.endswith("ail"):
if noun in _AIL_AUX:
return noun[:-3] + "aux", True # travail->travaux
return noun + "s", True
return noun + "s", True # default
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if number == "singular":
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
if d and d.get("PL"):
return d["PL"], "lexicon"
g = gender or noun_gender(lemma)
form, ok = _rule_plural(lemma, g)
return form, ("rule" if ok else "fallback")
# adjectives whose kaikki entries are unreliable: audited forms
_ADJ_FIX = {
"beau": {("m", "SG"): "beau", ("f", "SG"): "belle",
("m", "PL"): "beaux", ("f", "PL"): "belles"},
"nouveau": {("m", "SG"): "nouveau", ("f", "SG"): "nouvelle",
("m", "PL"): "nouveaux", ("f", "PL"): "nouvelles"},
"vieux": {("m", "SG"): "vieux", ("f", "SG"): "vieille",
("m", "PL"): "vieux", ("f", "PL"): "vieilles"},
"fou": {("m", "SG"): "fou", ("f", "SG"): "folle",
("m", "PL"): "fous", ("f", "PL"): "folles"},
"blanc": {("m", "SG"): "blanc", ("f", "SG"): "blanche",
("m", "PL"): "blancs", ("f", "PL"): "blanches"},
"long": {("m", "SG"): "long", ("f", "SG"): "longue",
("m", "PL"): "longs", ("f", "PL"): "longues"},
"bon": {("m", "SG"): "bon", ("f", "SG"): "bonne",
("m", "PL"): "bons", ("f", "PL"): "bonnes"},
}
def _rule_fem(a):
if a.endswith("e"):
return a
if a.endswith("er"):
return a[:-2] + "ère"
if a.endswith("eau"):
return a[:-3] + "elle"
if a.endswith("eux"):
return a[:-3] + "euse"
if a.endswith("f"):
return a[:-1] + "ve"
if a.endswith(("on", "en", "el", "eil", "et")):
return a + a[-1] + "e" # bon->bonne, ancien->ancienne, muet->muette
if a.endswith("c"):
return a[:-1] + "che" # blanc->blanche (public->publique via FIX)
return a + "e" # grand->grande, petit->petite, vert->verte
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
fix = _ADJ_FIX.get(lemma)
if fix and (g, num) in fix:
return fix[(g, num)], "lexicon"
d = _ADJS.get(lemma)
if d and d.get((g, num)):
return d[(g, num)], "lexicon"
# derive
msc = (d.get(("m", "SG")) if d else None) or lemma
if g == "m" and num == "SG":
return msc, "lexicon" if d else "rule"
fem = (d.get(("f", "SG")) if d else None) or _rule_fem(msc)
if g == "f" and num == "SG":
return fem, "lexicon" if (d and d.get(("f", "SG"))) else "rule"
if g == "m" and num == "PL":
if msc.endswith(("s", "x")):
return msc, "rule"
if msc.endswith("al"):
return msc[:-2] + "aux", "rule"
if msc.endswith("eau"):
return msc + "x", "rule"
return msc + "s", "rule"
# f|PL
return (fem if fem.endswith("s") else fem + "s"), "rule"
def lexicon_stats():
return {
"verb_source": "UniMorph French (github.com/unimorph/fra) + kaikki.org "
"irregulars (être + high-frequency)",
"noun_adj_source": "kaikki.org French (Wiktionary extract)",
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
"unimorph_verb_forms": len(_VERBS),
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
"irregular_verb_lemmas": len(_IRREGV),
"participle_lemmas": len(_PART),
"gerund_lemmas": len(_GER),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("parler", "ind", "present", "first", "singular", "parle"),
("être", "ind", "present", "third", "singular", "est"),
("avoir", "ind", "present", "first", "singular", "ai"),
("aller", "ind", "present", "third", "plural", "vont"),
("finir", "ind", "present", "first", "singular", "finis"),
("finir", "ind", "present", "first", "plural", "finissons"),
("manger", "ind", "present", "first", "plural", "mangeons"),
("faire", "ind", "future", "first", "singular", "ferai"),
("pouvoir", "sbjv", "present", "third", "singular", "puisse"),
("prendre", "ind", "passe_simple", "third", "singular", "prit"),
("vendre", "ind", "present", "third", "singular", "vend"),
("commencer", "ind", "imperfect", "first", "singular", "commençais"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
ok += got == exp
print(f" {flag}{lemma:10} {mood}/{tense:12} {per[:3]}.{num[:2]} -> {got:12} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" gender: maison=", noun_gender("maison"), "chat=", noun_gender("chat"),
"cheval=", noun_gender("cheval"), "nation=", noun_gender("nation"))
print(" plural: cheval->", inflect_noun("cheval", "plural"),
"| bateau->", inflect_noun("bateau", "plural"),
"| prix->", inflect_noun("prix", "plural"),
"| chat->", inflect_noun("chat", "plural"))
print(" adj: petit/f/sg->", inflect_adj("petit", "f", "singular"),
"| beau/f/sg->", inflect_adj("beau", "f", "singular"),
"| heureux/f/sg->", inflect_adj("heureux", "f", "singular"),
"| national/m/pl->", inflect_adj("national", "m", "plural"))
print(" part: aller/f/sg->", participle("aller", "f", "singular"),
"| prendre/f/pl->", participle("prendre", "f", "plural"),
"| finir/m/pl->", participle("finir", "m", "plural"))
print(" ger: manger->", gerund("manger"), "| finir->", gerund("finir"))
-588
View File
@@ -1,588 +0,0 @@
"""morphology_it_full.py — production-grade Italian morphological generator.
NOT a toy. Backed by three real, Wiktionary-lineage lexical sources:
VERBS
UniMorph Italian (github.com/unimorph/ita, CC-BY-SA 3.0)
10,009 verb lemmas × full paradigm, CLEAN orthography (no stress marks):
indicative present / imperfetto (PST;IPFV) / passato remoto (PST;PFV) /
futuro, condizionale (COND),
congiuntivo presente (SBJV;PRS) / imperfetto (SBJV;PST),
affirmative imperative, infinitive, gerundio (V.CVB;PRS),
past participle (masc-sg; fem/plural derived by vowel rule).
it_irreg_verbs.json — 66 high-frequency verbs UniMorph MISSES
(essere, avere, potere, uscire, tenere, prendere, piacere, …), extracted
from kaikki.org Italian, filtered to standard forms, and DE-STRESSED to
real orthography (kaikki marks tonic stress everywhere: pàrlo->parlo,
avùto->avuto; final legit accents kept: sarò, è). Built by build_it_irreg.py.
This layer takes priority — it supplies the two auxiliaries essere/avere,
which the whole passato-prossimo / essere-agreement system depends on.
NOUNS + ADJECTIVES — kaikki.org Italian (Wiktionary extract, CC-BY-SA 3.0)
noun lemmas WITH inherent gender (head-template arg) + real (often irregular)
plural — uomo->uomini, uovo->uova, dito->dita, città invariant — resolved
PER LEMMA, never guessed.
adjective lemmas with real feminine + masc/fem plural (italiano->italiana/
italiani/italiane, felice->felici invariant).
Fallbacks (degrade, never crash, on OOV input):
verbs : rule generator for regular -are/-ere/-ire (with -care/-gare h-insertion
and -ciare/-giare/-iare i-drop spelling rules)
nouns : gender heuristic (endings) + rule pluralization (ambiguous -co/-go FLAGGED)
adjs : -o/-a/-e gender rule + rule pluralization
Confidence flag on every form:
"lexicon" from UniMorph / kaikki-irregular / kaikki noun-adj (trust: high)
"rule" deterministic rule (trust: medium)
"fallback" could not inflect; returned lemma / ambiguous (trust: low -> FLAG)
Public API (used by realizer_it.py):
conjugate(lemma, mood, tense, person, number) -> (form, conf)
participle(lemma, gender="m", number="singular") -> (form, conf)
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"
inflect_noun(lemma, number, gender=None) -> (form, conf)
inflect_adj(lemma, gender, number) -> (form, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "ita.unimorph")
_IRREG = os.path.join(_HERE, "data", "it_irreg_verbs.json")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_it.jsonl")
_CACHE = os.path.join(_HERE, "data", "it_morph_cache.pkl")
# ── (mood, tense) -> UniMorph feature set that must ALL be present ────────────────
_VERB_KEYMAP = {
("ind", "present"): {"IND", "PRS"},
("ind", "imperfect"): {"IND", "PST", "IPFV"},
("ind", "passato_remoto"): {"IND", "PST", "PFV"},
("ind", "future"): {"IND", "FUT"},
("ind", "conditional"): {"COND"},
("sbjv", "present"): {"SBJV", "PRS"},
("sbjv", "imperfect"): {"SBJV", "PST"},
("imp", "affirmative"): {"POS", "IMP"},
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── build verb lexicon from UniMorph ─────────────────────────────────────────────
def _build_verbs():
verbs = {} # (lemma, "mood|tense|person|number") -> form
part = {} # lemma -> masc-sg past participle
ger = {} # lemma -> gerundio
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f:
part.setdefault(lemma, form)
continue
if head == "V.CVB": # gerundio (converb, present)
if "PRS" in f:
ger.setdefault(lemma, form)
continue
if head != "V":
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
for (mood, tense), req in _VERB_KEYMAP.items():
# exact-set discipline: PST;PFV must not match PST;IPFV, etc.
if not req <= f:
continue
# guard IND;PST ambiguity: require the specific aspect feature
if tense == "imperfect" and "PFV" in f:
continue
if tense == "passato_remoto" and "IPFV" in f:
continue
# COND must not also be a subjunctive/imperative slot
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), form)
break
return verbs, part, ger
# ── kaikki nouns + adjectives ────────────────────────────────────────────────────
_EXCL_FORM_TAGS = {"alternative", "archaic", "obsolete", "dialectal", "regional",
"diminutive", "augmentative", "pejorative", "comparative",
"superlative", "misspelling", "rare", "informal", "literary",
"poetic", "error-unrecognized-form", "apocopic", "obsolete",
"construed", "collective"}
def _kaikki_gender(arg):
if not arg:
return None
a = str(arg).lower()
if a.startswith("f"):
return "f"
if a.startswith("m"):
return "m"
return None
def _build_nouns_adjs():
nouns = {} # lemma -> {"g","SG","PL"}
adjs = {} # lemma -> {("m","SG"),("f","SG"),("m","PL"),("f","PL")}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
pos = d.get("pos")
word = d.get("word", "")
if not word or " " in word:
continue
forms = d.get("forms", []) or []
if pos == "noun":
ht = d.get("head_templates") or []
g = None
if ht:
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
if g is None:
tags = d.get("tags") or []
if "feminine" in tags:
g = "f"
elif "masculine" in tags:
g = "m"
pl = None
for x in forms:
t = set(x.get("tags") or [])
if "plural" in t and not (t & _EXCL_FORM_TAGS):
fm = x.get("form")
if fm and " " not in fm and fm != "#":
pl = fm
break
if word not in nouns:
nouns[word] = {"g": g, "SG": word, "PL": pl}
else:
cur = nouns[word]
if cur.get("g") is None and g:
cur["g"] = g
if not cur.get("PL") and pl:
cur["PL"] = pl
elif pos == "adj":
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word)
for x in forms:
t = set(x.get("tags") or [])
fm = x.get("form")
if not fm or " " in fm or (t & _EXCL_FORM_TAGS):
continue
if "feminine" in t and "plural" in t:
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
elif "masculine" in t and "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
elif "feminine" in t:
d0[("f", "SG")] = d0.get(("f", "SG")) or fm
elif "plural" in t: # invariant-gender adj (felice -> felici)
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
return nouns, adjs
def _build_cache():
verbs, part, ger = _build_verbs()
nouns, adjs = _build_nouns_adjs()
with open(_IRREG, encoding="utf-8") as fh:
irreg = json.load(fh)
data = {"verbs": verbs, "part": part, "ger": ger,
"nouns": nouns, "adjs": adjs, "irreg": irreg}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [_UNIMORPH, _KAIKKI, _IRREG]
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PART, _GER, _NOUNS, _ADJS, _IRREGV = (
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"],
_LEX["irreg"])
# ── regular-ending rule fallback ─────────────────────────────────────────────────
def _vclass(lemma):
if lemma.endswith("are"):
return "are"
if lemma.endswith("ere"):
return "ere"
if lemma.endswith("ire"):
return "ire"
return None
# endings [1sg,2sg,3sg,1pl,2pl,3pl]
_REG = {
("ind", "present", "are"): ["o", "i", "a", "iamo", "ate", "ano"],
("ind", "present", "ere"): ["o", "i", "e", "iamo", "ete", "ono"],
("ind", "present", "ire"): ["o", "i", "e", "iamo", "ite", "ono"],
("ind", "imperfect", "are"): ["avo", "avi", "ava", "avamo", "avate", "avano"],
("ind", "imperfect", "ere"): ["evo", "evi", "eva", "evamo", "evate", "evano"],
("ind", "imperfect", "ire"): ["ivo", "ivi", "iva", "ivamo", "ivate", "ivano"],
("ind", "passato_remoto", "are"): ["ai", "asti", "ò", "ammo", "aste", "arono"],
("ind", "passato_remoto", "ere"): ["ei", "esti", "é", "emmo", "este", "erono"],
("ind", "passato_remoto", "ire"): ["ii", "isti", "ì", "immo", "iste", "irono"],
("sbjv", "present", "are"): ["i", "i", "i", "iamo", "iate", "ino"],
("sbjv", "present", "ere"): ["a", "a", "a", "iamo", "iate", "ano"],
("sbjv", "present", "ire"): ["a", "a", "a", "iamo", "iate", "ano"],
("sbjv", "imperfect", "are"): ["assi", "assi", "asse", "assimo", "aste", "assero"],
("sbjv", "imperfect", "ere"): ["essi", "essi", "esse", "essimo", "este", "essero"],
("sbjv", "imperfect", "ire"): ["issi", "issi", "isse", "issimo", "iste", "issero"],
# imperative: 2sg,3sg(Lei),1pl,2pl,3pl (1sg has none)
("imp", "affirmative", "are"): [None, "a", "i", "iamo", "ate", "ino"],
("imp", "affirmative", "ere"): [None, "i", "a", "iamo", "ete", "ano"],
("imp", "affirmative", "ire"): [None, "i", "a", "iamo", "ite", "ano"],
}
# future / conditional attach to a stem = infinitive minus final -e, with
# -are -> -er (parlare->parler-), -ere/-ire keep (credere->creder-, dormir-)
_FUT = ["ò", "ai", "à", "emo", "ete", "anno"]
_COND = ["ei", "esti", "ebbe", "emmo", "este", "ebbero"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _fut_stem(lemma, vc):
body = lemma[:-3] # drop are/ere/ire
if vc == "are":
return body + "er"
return body + vc[0] + "r" # ere->er? no: keep vowel: creder-, dormir-
# NOTE corrected below
def _apply_are_spelling(stem, ending):
"""-care/-gare insert h before front endings; -ciare/-giare/-sciare/-iare drop i."""
front = ending[:1] in ("i", "e")
if stem.endswith(("c", "g")) and front:
return stem + "h" + ending
if stem.endswith(("ci", "gi", "sci")) and ending[:1] == "i":
return stem[:-1] + ending # mangi+iamo -> mangiamo
if stem.endswith("i") and ending[:1] == "i":
return stem[:-1] + ending # studi+iamo -> studiamo
return stem + ending
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
body = lemma[:-3]
i = _slot_idx(person, number)
if mood == "ind" and tense in ("future", "conditional"):
stem = body + "er" if vc == "are" else body + vc[0] + "r"
# ere: creder-, ire: dormir- -> body + 'e'/'i' + 'r'
if vc == "ere":
stem = body + "er"
elif vc == "ire":
stem = body + "ir"
end = (_FUT if tense == "future" else _COND)[i]
# spelling: -care/-gare -> cherò/gherò ; -ciare/-giare -> cerò/gerò
if vc == "are":
if body.endswith(("c", "g")):
stem = body + "her"
elif body.endswith(("ci", "gi", "sci")):
stem = body[:-1] + "er"
elif body.endswith("i"):
stem = body[:-1] + "er"
return stem + end
table = _REG.get((mood, tense, vc))
if not table:
return None
end = table[i]
if end is None:
return None
if vc == "are":
return _apply_are_spelling(body, end)
# -ere/-ire: guard against double-i (dormi+iamo -> dormiamo)
if body.endswith("i") and end[:1] == "i":
return body[:-1] + end
return body + end
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
"""Return (surface, confidence). mood in ind|sbjv|imp; tense per _VERB_KEYMAP."""
lemma = lemma.strip().lower()
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{number}"
ir = _IRREGV.get(lemma)
if ir and key in ir:
return ir[key], "lexicon"
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
if form:
return form, "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r:
return r, "rule"
return lemma, "fallback"
# ── PUBLIC: participle + gerund ──────────────────────────────────────────────────
def _participle_msg(lemma):
"""Return (masc-sg participle, source) or (None, None)."""
ir = _IRREGV.get(lemma)
if ir and "part" in ir:
return ir["part"], "lexicon"
if lemma in _PART:
return _PART[lemma], "lexicon"
return None, None
def participle(lemma, gender="m", number="singular"):
"""Past participle with gender/number agreement (for essere-perfect & passives).
UniMorph/irregular give masc-sg; fem/plural derived by final-vowel swap
(-o -> -a/-i/-e), valid for regular -ato/-uto/-ito AND irregulars
(preso->presa/presi/prese, aperto->aperta/aperti/aperte, morto->morta/...)."""
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
msg, src = _participle_msg(lemma)
conf = "lexicon"
if msg is None:
vc = _vclass(lemma)
if vc == "are":
msg = lemma[:-3] + "ato"
elif vc == "ere":
msg = lemma[:-3] + "uto"
elif vc == "ire":
msg = lemma[:-3] + "ito"
else:
return lemma, "fallback"
conf = "rule"
# agreement: only -o participles inflect for gender+number
if msg.endswith("o"):
stem = msg[:-1]
suf = {"m|SG": "o", "f|SG": "a", "m|PL": "i", "f|PL": "e"}[f"{g}|{num}"]
return stem + suf, conf
return msg, conf # non -o participle: leave as-is (rare)
def gerund(lemma):
lemma = lemma.strip().lower()
ir = _IRREGV.get(lemma)
if ir and "ger" in ir:
return ir["ger"], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
vc = _vclass(lemma)
if vc == "are":
return lemma[:-3] + "ando", "rule"
if vc in ("ere", "ire"):
return lemma[:-3] + "endo", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
_FEM_SUF = ("zione", "sione", "gione", "", "", "trice", "aggine", "udine",
"igine", "ie", "essa", "izia", "ezza")
_MASC_SUF = ("ore", "ame", "iere", "ale", "ile")
def _gender_heuristic(noun):
for suf in _FEM_SUF:
if noun.endswith(suf):
return "f"
for suf in _MASC_SUF:
if noun.endswith(suf):
return "m"
if noun.endswith("o"):
return "m"
if noun.endswith("a"):
return "f"
if noun.endswith("à") or noun.endswith("ù"):
return "f"
return "m" # -e and consonant-final loanwords default masculine
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g") in ("m", "f"):
return d["g"]
return _gender_heuristic(lemma)
def _rule_plural(noun, gender):
"""Deterministic Italian pluralization. Returns (form, ok); ok=False FLAGS an
ambiguous case the lexicon would normally resolve (-co/-go palatalization)."""
if not noun:
return noun, True
# invariant: accented final vowel, consonant-final, monosyllable, -i final
if noun[-1:] in ("à", "è", "é", "ì", "í", "ò", "ó", "ù", "ú"):
return noun, True
if noun[-1:] not in ("a", "e", "o", "i", "u"):
return noun, True # consonant-final loanword: invariant
if noun.endswith("i"):
return noun, True # e.g. crisi, analisi: invariant
if noun.endswith("io"):
return noun[:-2] + "i", True # figlio->figli (unstressed i)
if noun.endswith("cia") or noun.endswith("gia"):
# vowel before cia/gia -> -cie/-gie ; consonant -> -ce/-ge (approx)
return noun[:-2] + "e", True # arancia->arance (majority)
if noun.endswith("ca"):
return noun[:-2] + "che", True # amica->amiche
if noun.endswith("ga"):
return noun[:-2] + "ghe", True
if noun.endswith("co"):
return noun[:-2] + "chi", False # AMBIGUOUS (amico->amici) -> flag
if noun.endswith("go"):
return noun[:-2] + "ghi", False # AMBIGUOUS (psicologo->psicologi)
if noun.endswith("a"):
return noun[:-1] + "e", True # casa->case (m -a: -i, but rare)
if noun.endswith("o"):
return noun[:-1] + "i", True # libro->libri
if noun.endswith("e"):
return noun[:-1] + "i", True # cane->cani, chiave->chiavi
return noun, True
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if number == "singular":
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
if d and d.get("PL"):
return d["PL"], "lexicon"
g = gender or noun_gender(lemma)
form, ok = _rule_plural(lemma, g)
return form, ("rule" if ok else "fallback")
# adjectives whose kaikki entries are unreliable (messy inflection templates):
# supply audited regular agreement forms (prenominal apocope handled in realizer).
_ADJ_FIX = {
"bello": {("m", "SG"): "bello", ("f", "SG"): "bella",
("m", "PL"): "belli", ("f", "PL"): "belle"},
"quello": {("m", "SG"): "quello", ("f", "SG"): "quella",
("m", "PL"): "quelli", ("f", "PL"): "quelle"},
}
# ── PUBLIC: adjective agreement ──────────────────────────────────────────────────
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
fix = _ADJ_FIX.get(lemma)
if fix and (g, num) in fix:
return fix[(g, num)], "lexicon"
d = _ADJS.get(lemma)
if d:
form = d.get((g, num))
if form:
return form, "lexicon"
sg = d.get((g, "SG")) or d.get(("m", "SG")) or lemma
if num == "PL":
pl, ok = _rule_plural(sg, g)
return pl, ("rule" if ok else "fallback")
return sg, "lexicon"
# rule fallback
a = lemma
if a.endswith("o"): # -o/-a/-i/-e class
base = a[:-1]
suf = {"m|SG": "o", "f|SG": "a", "m|PL": "i", "f|PL": "e"}[f"{g}|{num}"]
return base + suf, "rule"
if a.endswith("e"): # felice-class: SG invariant, PL -i
if num == "PL":
return a[:-1] + "i", "rule"
return a, "rule"
if num == "PL":
p, ok = _rule_plural(a, g)
return p, ("rule" if ok else "fallback")
return a, "rule"
def lexicon_stats():
return {
"verb_source": "UniMorph Italian (github.com/unimorph/ita) + kaikki.org "
"irregulars (de-stressed)",
"noun_adj_source": "kaikki.org Italian (Wiktionary extract)",
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
"unimorph_verb_forms": len(_VERBS),
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
"irregular_verb_lemmas": len(_IRREGV),
"participle_lemmas": len(_PART),
"gerund_lemmas": len(_GER),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("parlare", "ind", "present", "first", "singular", "parlo"),
("essere", "ind", "present", "third", "singular", "è"),
("avere", "ind", "present", "first", "singular", "ho"),
("mangiare", "ind", "present", "second", "singular", "mangi"),
("finire", "ind", "present", "first", "singular", "finisco"),
("andare", "ind", "present", "third", "plural", "vanno"),
("fare", "ind", "future", "first", "singular", "farò"),
("potere", "sbjv", "present", "third", "singular", "possa"),
("prendere", "ind", "passato_remoto", "first", "singular", "presi"),
("cercare", "ind", "present", "second", "singular", "cerchi"),
("dormire", "ind", "present", "third", "plural", "dormono"),
("credere", "ind", "future", "first", "singular", "crederò"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
ok += got == exp
print(f" {flag}{lemma:9} {mood}/{tense:14} {per[:3]}.{num[:2]} -> {got:12} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" gender: casa=", noun_gender("casa"), "problema=", noun_gender("problema"),
"mano=", noun_gender("mano"), "città=", noun_gender("città"),
"cane=", noun_gender("cane"))
print(" plural: uomo->", inflect_noun("uomo", "plural"),
"| uovo->", inflect_noun("uovo", "plural"),
"| città->", inflect_noun("città", "plural"),
"| amico->", inflect_noun("amico", "plural"),
"| casa->", inflect_noun("casa", "plural"))
print(" adj: italiano/f/pl->", inflect_adj("italiano", "f", "plural"),
"| felice/m/pl->", inflect_adj("felice", "m", "plural"),
"| bello/f/sg->", inflect_adj("bello", "f", "singular"))
print(" part: aprire/f/sg->", participle("aprire", "f", "singular"),
"| prendere/m/pl->", participle("prendere", "m", "plural"),
"| andare/f/sg->", participle("andare", "f", "singular"))
print(" ger: fare->", gerund("fare"), "| parlare->", gerund("parlare"))
-666
View File
@@ -1,666 +0,0 @@
# -*- coding: utf-8 -*-
"""morphology_lat_full.py — production-grade Latin morphological generator.
Latin is the FLAGSHIP dead-language realizer. It rides the *architecture* of the
Romance/Italic engine (the same Realization / spec-driven design and the UniMorph
loader pattern from morphology_it_full.py) but with the CASE SYSTEM RESTORED
the feature Romance lost. Latin therefore exercises machinery the modern Romance
siblings never needed: 5 declensions x 6 cases x 2 numbers x 3 genders, plus a
4-conjugation verb system with tense/mood/voice.
DATA (real, attested no fabrication):
NOUNS + ADJECTIVES UniMorph Latin (github.com/unimorph/lat, CC-BY-SA 3.0)
163,182 N forms across ~thousands of lemmas, each with the full case paradigm
N;NOM/GEN/DAT/ACC/ABL/VOC;SG/PL (real inflected forms, WITH macrons:
puella->puellam, rēx->rēgis, corpus->corporis).
244,197 ADJ forms with case x GENDER x number, incl. UniMorph's combined
tags (GEN+DAT, MASC+FEM, MASC+FEM+NEUT) which are split on load.
462,668 V.PTCP forms (participles) also carry case/gender/number.
UniMorph N tags DO NOT encode inherent gender, so noun gender is inferred
from the declension (nom-sg + gen-sg endings) with a curated exceptions
map the standard, attestable rule (1st decl -a/-ae = fem, 2nd -us/-i =
masc, -um = neut, ...).
VERBS RULE ENGINE (honest gap: UniMorph Latin's verb list is a 947-lemma
sample of rare/prefixed verbs that MISSES every core textbook verb amō,
videō, sum, regō, ... are all absent). Latin conjugation is, however, highly
regular, so verbs are generated by a deterministic 4-conjugation engine over
curated principal parts (present / perfect / supine stems), sourced from
standard references. Irregulars (sum, possum, , ferō, volō, nōlō, mālō)
are curated full tables. Forms are flagged "rule" (not "lexicon") for honesty.
Confidence flag on every form (same contract as the Romance engine):
"lexicon" from UniMorph (trust: high)
"rule" deterministic morphology rule (trust: medium)
"fallback" could not inflect; returned lemma (trust: low -> FLAG)
Public API (used by realizer_lat.py):
decline_noun(lemma, case, number) -> (form, conf)
noun_gender(lemma) -> "m"|"f"|"n"
decline_adj(lemma, case, gender, number) -> (form, conf)
conjugate(lemma, tense, mood, voice, person, number) -> (form, conf)
participle(lemma, kind, case, gender, number) -> (form, conf) # kind: prs|pfv|fut
infinitive(lemma, tense="present", voice="active") -> (form, conf)
lexicon_stats() -> dict
"""
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "lat.unimorph")
_CACHE = os.path.join(_HERE, "data", "lat_morph_cache.pkl")
_CASES = ("NOM", "GEN", "DAT", "ACC", "ABL", "VOC")
_CASE_MAP = {"nom": "NOM", "gen": "GEN", "dat": "DAT", "acc": "ACC",
"abl": "ABL", "voc": "VOC"}
_NUM = {"singular": "SG", "plural": "PL"}
_GEN = {"m": "MASC", "f": "FEM", "n": "NEUT"}
# ── UniMorph loader: noun + adjective + participle case paradigms ────────────────
def _build_cache():
nouns = {} # lemma -> {(CASE, NUM): form}
adjs = {} # lemma -> {(CASE, GEN, NUM): form}
ptcps = {} # lemma -> {(CASE, GEN, NUM): form} (from V.PTCP; keyed loosely)
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
feats = tag.split(";")
head = feats[0]
fs = set(feats)
case = next((c for c in _CASES if c in fs), None)
# handle combined case tags like GEN+DAT
if case is None:
for f in feats:
if "+" in f and any(c in f.split("+") for c in _CASES):
case = [c for c in _CASES if c in f.split("+")]
break
num = "SG" if "SG" in fs else ("PL" if "PL" in fs else None)
if case is None or num is None:
continue
cases = case if isinstance(case, list) else [case]
if head == "N":
d = nouns.setdefault(lemma, {})
for c in cases:
d.setdefault((c, num), form)
elif head == "ADJ":
# gender may be combined: MASC+FEM+NEUT, MASC+FEM
genders = []
for g in ("MASC", "FEM", "NEUT"):
if any(g == x or (g in x.split("+")) for x in feats):
genders.append(g)
if not genders:
genders = ["MASC", "FEM", "NEUT"]
d = adjs.setdefault(lemma, {})
for c in cases:
for g in genders:
d.setdefault((c, g, num), form)
data = {"nouns": nouns, "adjs": adjs, "ptcps": ptcps}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE) and os.path.exists(_UNIMORPH):
if os.path.getmtime(_CACHE) >= os.path.getmtime(_UNIMORPH):
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_NOUNS, _ADJS = _LEX["nouns"], _LEX["adjs"]
# ── noun gender inference (declension-based, curated exceptions) ─────────────────
# Real, attestable rule: gender follows declension + nominative shape, with the
# standard closed set of exceptions.
_GENDER_EXC = {
# 1st-declension masculines (people/agents)
"agricola": "m", "poēta": "m", "nauta": "m", "incola": "m", "scrība": "m",
"auriga": "m", "pīrāta": "m", "athlēta": "m",
# 2nd-declension neuters / feminines
"vīrus": "n", "vulgus": "n", "pelagus": "n", "humus": "f",
# common 3rd-declension whose gender the ending would mispredict
"rēx": "m", "dux": "m", "mīles": "m", "pater": "m", "frāter": "m",
"homō": "m", "leō": "m", "sōl": "m", "mōns": "m", "pōns": "m", "fōns": "m",
"sanguis": "m", "ōrdō": "m", "sermō": "m", "amor": "m", "dolor": "m",
"labor": "m", "timor": "m", "honor": "m", "color": "m", "pēs": "m",
"dēns": "m", "flōs": "m", "mōs": "m", "mensis": "m", "orbis": "m",
"piscis": "m", "ignis": "m", "collis": "m", "grex": "m", "prīnceps": "m",
"māter": "f", "soror": "f", "uxor": "f", "mulier": "f", "virgō": "f",
"urbs": "f", "arx": "f", "pāx": "f", "lēx": "f", "lūx": "f", "vōx": "f",
"nox": "f", "nix": "f", "vīs": "f", "salūs": "f", "virtūs": "f",
"aetās": "f", "cīvitās": "f", "lībertās": "f", "vēritās": "f", "voluptās": "f",
"nātiō": "f", "ratiō": "f", "ōrātiō": "f", "legiō": "f", "regiō": "f",
"mens": "f", "gens": "f", "ars": "f", "pars": "f", "mors": "f", "sors": "f",
"nāvis": "f", "turris": "f", "avis": "f", "vallis": "f", "classis": "f",
"corpus": "n", "tempus": "n", "opus": "n", "genus": "n", "onus": "n",
"pectus": "n", "latus": "n", "vulnus": "n", "scelus": "n", "sīdus": "n",
"caput": "n", "iter": "n", "flūmen": "n", "nōmen": "n", "carmen": "n",
"agmen": "n", "certāmen": "n", "lūmen": "n", "ōmen": "n", "cōgnōmen": "n",
"mare": "n", "animal": "n", "exemplar": "n", "rēte": "n",
# 4th-declension exceptions
"manus": "f", "domus": "f", "tribus": "f", "porticus": "f", "īdūs": "f",
"cornū": "n", "genū": "n", "gelū": "n", "verū": "n",
# 5th-declension
"diēs": "m", "merīdiēs": "m",
}
def _infer_gender(lemma):
if lemma in _GENDER_EXC:
return _GENDER_EXC[lemma]
d = _NOUNS.get(lemma)
nom = d.get(("NOM", "SG")) if d else lemma
gen = d.get(("GEN", "SG")) if d else None
nom = nom or lemma
# 5th declension: gen -eī / -ēī
if gen and (gen.endswith("") or gen.endswith("ēī")):
return "f"
# 1st declension: nom -a, gen -ae
if nom.endswith("a") and (not gen or gen.endswith("ae")):
return "f"
# 2nd declension neuter: nom -um
if nom.endswith("um"):
return "n"
# 2nd declension masc: nom -us/-er/-ir, gen -ī
if (nom.endswith("us") or nom.endswith("er") or nom.endswith("ir")) and \
(not gen or gen.endswith("ī")):
return "m"
# 4th declension: gen -ūs
if gen and gen.endswith("ūs"):
return "n" if nom.endswith("ū") else "m"
# 3rd declension neuters by common nom endings
if nom.endswith(("men", "us", "ur", "al", "ar", "e", "ma")):
# -us here is 3rd-decl neuter type (corpus) only if gen shows -oris/-eris
if nom.endswith("us") and gen and (gen.endswith("oris") or gen.endswith("eris")
or gen.endswith("uris")):
return "n"
if nom.endswith(("men", "al", "ar", "e")):
return "n"
# default 3rd-declension: masculine (most common)
return "m"
_GENDER_CACHE = {}
def noun_gender(lemma):
lemma = lemma.strip()
if lemma not in _GENDER_CACHE:
_GENDER_CACHE[lemma] = _infer_gender(lemma)
return _GENDER_CACHE[lemma]
# ── PUBLIC: noun declension ─────────────────────────────────────────────────────
def decline_noun(lemma, case, number):
lemma = lemma.strip()
C = _CASE_MAP.get(case, case.upper())
N = _NUM.get(number, number)
d = _NOUNS.get(lemma)
if d and (C, N) in d:
return d[(C, N)], "lexicon"
# abl sg often == the -e/-o form; try nom fallback
if d:
# try VOC==NOM, ACC neuter==NOM etc are already in data; last resort lemma
return lemma, "fallback"
return lemma, "fallback"
# ── PUBLIC: adjective declension ────────────────────────────────────────────────
def decline_adj(lemma, case, gender, number):
lemma = lemma.strip()
C = _CASE_MAP.get(case, case.upper())
G = _GEN.get(gender, gender.upper())
N = _NUM.get(number, number)
d = _ADJS.get(lemma)
if d and (C, G, N) in d:
return d[(C, G, N)], "lexicon"
# try other gender (some adjs listed only under MASC+FEM etc handled at load)
if d:
for altG in ("MASC", "FEM", "NEUT"):
if (C, altG, N) in d:
return d[(C, altG, N)], "lexicon"
return lemma, "fallback"
return lemma, "fallback"
# ═══════════════════════════════════════════════════════════════════════════════
# VERB RULE ENGINE (4 conjugations + curated irregulars)
# ═══════════════════════════════════════════════════════════════════════════════
# Curated principal parts for common attested verbs:
# lemma -> (conj, present_stem, perfect_stem, supine_stem)
# conj in {1,2,3,"3io",4}. Stems carry macrons (matching UniMorph orthography).
_VERBS = {
"amō": (1, "am", "amāv", "amāt"),
"laudō": (1, "laud", "laudāv", "laudāt"),
"portō": (1, "port", "portāv", "portāt"),
"vocō": (1, "voc", "vocāv", "vocāt"),
"": (1, "d", "ded", "dat"),
"spectō": (1, "spect", "spectāv", "spectāt"),
"pugnō": (1, "pugn", "pugnāv", "pugnāt"),
"labōrō": (1, "labōr", "labōrāv", "labōrāt"),
"necō": (1, "nec", "necāv", "necāt"),
"parō": (1, "par", "parāv", "parāt"),
"cōgitō": (1, "cōgit", "cōgitāv", "cōgitāt"),
"habitō": (1, "habit", "habitāv", "habitāt"),
"nārrō": (1, "nārr", "nārrāv", "nārrāt"),
"servō": (1, "serv", "servāv", "servāt"),
"superō": (1, "super", "superāv", "superāt"),
"oppugnō": (1, "oppugn", "oppugnāv", "oppugnāt"),
"ambulō": (1, "ambul", "ambulāv", "ambulāt"),
"clāmō": (1, "clām", "clāmāv", "clāmāt"),
"vulnerō": (1, "vulner", "vulnerāv", "vulnerāt"),
"aedificō": (1, "aedific", "aedificāv", "aedificāt"),
"expugnō": (1, "expugn", "expugnāv", "expugnāt"),
"dēfendō": (3, "dēfend", "dēfend", "dēfēns"),
"petō": (3, "pet", "petīv", "petīt"),
"occīdō": (3, "occīd", "occīd", "occīs"),
"interficiō": ("3io", "interfic", "interfēc", "interfect"),
"timeō": (2, "tim", "timu", None),
"iaceō": (2, "iac", "iacu", None),
"pāreō": (2, "pār", "pāru", "pārit"),
"respondeō": (2, "respond", "respond", "respōns"),
"vertō": (3, "vert", "vert", "vers"),
"ostendō": (3, "ostend", "ostend", "ostent"),
"cōnstituō": (3, "cōnstitu", "cōnstitu", "cōnstitūt"),
"cōgnōscō": (3, "cōgnōsc", "cōgnōv", "cōgnit"),
"crēdō": (3, "crēd", "crēdid", "crēdit"),
"ēdūcō": (3, "ēdūc", "ēdūx", "ēduct"),
"cōnservō": (1, "cōnserv", "cōnservāv", "cōnservāt"),
"iuvō": (1, "iuv", "iūv", "iūt"),
"dēbeō": (2, "dēb", "dēbu", "dēbit"),
"moneō": (2, "mon", "monu", "monit"),
"videō": (2, "vid", "vīd", "vīs"),
"habeō": (2, "hab", "habu", "habit"),
"teneō": (2, "ten", "tenu", "tent"),
"timeō": (2, "tim", "timu", None),
"terreō": (2, "terr", "terru", "territ"),
"dēleō": (2, "dēl", "dēlēv", "dēlēt"),
"iubeō": (2, "iub", "iuss", "iuss"),
"maneō": (2, "man", "māns", "māns"),
"moveō": (2, "mov", "mōv", "mōt"),
"doceō": (2, "doc", "docu", "doct"),
"sedeō": (2, "sed", "sēd", "sess"),
"rīdeō": (2, "rīd", "rīs", "rīs"),
"regō": (3, "reg", "rēx", "rēct"),
"dūcō": (3, "dūc", "dūx", "duct"),
"scrībō": (3, "scrīb", "scrīps", "scrīpt"),
"mittō": (3, "mitt", "mīs", "miss"),
"pōnō": (3, "pōn", "posu", "posit"),
"agō": (3, "ag", "ēg", "āct"),
"dīcō": (3, "dīc", "dīx", "dict"),
"gerō": (3, "ger", "gess", "gest"),
"vincō": (3, "vinc", "vīc", "vict"),
"petō": (3, "pet", "petīv", "petīt"),
"legō": (3, "leg", "lēg", "lēct"),
"currō": (3, "curr", "cucurr", "curs"),
"vīvō": (3, "vīv", "vīx", "vīct"),
"quaerō": (3, "quaer", "quaesīv", "quaesīt"),
"trahō": (3, "trah", "trāx", "tract"),
"claudō": (3, "claud", "claus", "claus"),
"cōgō": (3, "cōg", "coēg", "coāct"),
"relinquō": (3, "relinqu", "relīqu", "relict"),
"capiō": ("3io", "cap", "cēp", "capt"),
"faciō": ("3io", "fac", "fēc", "fact"),
"iaciō": ("3io", "iac", "iēc", "iact"),
"rapiō": ("3io", "rap", "rapu", "rapt"),
"fugiō": ("3io", "fug", "fūg", "fugit"),
"cupiō": ("3io", "cup", "cupīv", "cupīt"),
"accipiō": ("3io", "accip", "accēp", "accept"),
"audiō": (4, "aud", "audīv", "audīt"),
"veniō": (4, "ven", "vēn", "vent"),
"sciō": (4, "sc", "scīv", "scīt"),
"sentiō": (4, "sent", "sēns", "sēns"),
"mūniō": (4, "mūn", "mūnīv", "mūnīt"),
"dormiō": (4, "dorm", "dormīv", "dormīt"),
"aperiō": (4, "aper", "aperu", "apert"),
"inveniō": (4, "inven", "invēn", "invent"),
}
# ── Present-system paradigms: full ending tables per conjugation, attached to the
# bare present stem (pstem). Hardcoded from the standard grammar with correct
# macrons/vowel-lengths — deterministic and independently verifiable. Keys:
# (tense, mood, voice) -> {conj: [1sg,2sg,3sg,1pl,2pl,3pl]}
_PARADIGM = {
("present", "ind", "active"): {
1: ["ō", "ās", "at", "āmus", "ātis", "ant"],
2: ["", "ēs", "et", "ēmus", "ētis", "ent"],
3: ["ō", "is", "it", "imus", "itis", "unt"],
"3io": ["", "is", "it", "imus", "itis", "iunt"],
4: ["", "īs", "it", "īmus", "ītis", "iunt"],
},
("present", "ind", "passive"): {
1: ["or", "āris", "ātur", "āmur", "āminī", "antur"],
2: ["eor", "ēris", "ētur", "ēmur", "ēminī", "entur"],
3: ["or", "eris", "itur", "imur", "iminī", "untur"],
"3io": ["ior", "eris", "itur", "imur", "iminī", "iuntur"],
4: ["ior", "īris", "ītur", "īmur", "īminī", "iuntur"],
},
("imperfect", "ind", "active"): {
1: ["ābam", "ābās", "ābat", "ābāmus", "ābātis", "ābant"],
2: ["ēbam", "ēbās", "ēbat", "ēbāmus", "ēbātis", "ēbant"],
3: ["ēbam", "ēbās", "ēbat", "ēbāmus", "ēbātis", "ēbant"],
"3io": ["iēbam", "iēbās", "iēbat", "iēbāmus", "iēbātis", "iēbant"],
4: ["iēbam", "iēbās", "iēbat", "iēbāmus", "iēbātis", "iēbant"],
},
("imperfect", "ind", "passive"): {
1: ["ābar", "ābāris", "ābātur", "ābāmur", "ābāminī", "ābantur"],
2: ["ēbar", "ēbāris", "ēbātur", "ēbāmur", "ēbāminī", "ēbantur"],
3: ["ēbar", "ēbāris", "ēbātur", "ēbāmur", "ēbāminī", "ēbantur"],
"3io": ["iēbar", "iēbāris", "iēbātur", "iēbāmur", "iēbāminī", "iēbantur"],
4: ["iēbar", "iēbāris", "iēbātur", "iēbāmur", "iēbāminī", "iēbantur"],
},
("future", "ind", "active"): {
1: ["ābō", "ābis", "ābit", "ābimus", "ābitis", "ābunt"],
2: ["ēbō", "ēbis", "ēbit", "ēbimus", "ēbitis", "ēbunt"],
3: ["am", "ēs", "et", "ēmus", "ētis", "ent"],
"3io": ["iam", "iēs", "iet", "iēmus", "iētis", "ient"],
4: ["iam", "iēs", "iet", "iēmus", "iētis", "ient"],
},
("future", "ind", "passive"): {
1: ["ābor", "āberis", "ābitur", "ābimur", "ābiminī", "ābuntur"],
2: ["ēbor", "ēberis", "ēbitur", "ēbimur", "ēbiminī", "ēbuntur"],
3: ["ar", "ēris", "ētur", "ēmur", "ēminī", "entur"],
"3io": ["iar", "iēris", "iētur", "iēmur", "iēminī", "ientur"],
4: ["iar", "iēris", "iētur", "iēmur", "iēminī", "ientur"],
},
("present", "sbjv", "active"): {
1: ["em", "ēs", "et", "ēmus", "ētis", "ent"],
2: ["eam", "eās", "eat", "eāmus", "eātis", "eant"],
3: ["am", "ās", "at", "āmus", "ātis", "ant"],
"3io": ["iam", "iās", "iat", "iāmus", "iātis", "iant"],
4: ["iam", "iās", "iat", "iāmus", "iātis", "iant"],
},
("present", "sbjv", "passive"): {
1: ["er", "ēris", "ētur", "ēmur", "ēminī", "entur"],
2: ["ear", "eāris", "eātur", "eāmur", "eāminī", "eantur"],
3: ["ar", "āris", "ātur", "āmur", "āminī", "antur"],
"3io": ["iar", "iāris", "iātur", "iāmur", "iāminī", "iantur"],
4: ["iar", "iāris", "iātur", "iāmur", "iāminī", "iantur"],
},
("imperfect", "sbjv", "active"): {
1: ["ārem", "ārēs", "āret", "ārēmus", "ārētis", "ārent"],
2: ["ērem", "ērēs", "ēret", "ērēmus", "ērētis", "ērent"],
3: ["erem", "erēs", "eret", "erēmus", "erētis", "erent"],
"3io": ["erem", "erēs", "eret", "erēmus", "erētis", "erent"],
4: ["īrem", "īrēs", "īret", "īrēmus", "īrētis", "īrent"],
},
("imperfect", "sbjv", "passive"): {
1: ["ārer", "ārēris", "ārētur", "ārēmur", "ārēminī", "ārentur"],
2: ["ērer", "ērēris", "ērētur", "ērēmur", "ērēminī", "ērentur"],
3: ["erer", "erēris", "erētur", "erēmur", "erēminī", "erentur"],
"3io": ["erer", "erēris", "erētur", "erēmur", "erēminī", "erentur"],
4: ["īrer", "īrēris", "īrētur", "īrēmur", "īrēminī", "īrentur"],
},
}
# perfect-active endings (added to perfect stem) — same for all conjugations
_PERF_ACT = {
("perfect", "ind"): ["ī", "istī", "it", "imus", "istis", "ērunt"],
("pluperfect", "ind"): ["eram", "erās", "erat", "erāmus", "erātis", "erant"],
("futureperfect", "ind"): ["erō", "eris", "erit", "erimus", "eritis", "erint"],
("perfect", "sbjv"): ["erim", "erīs", "erit", "erīmus", "erītis", "erint"],
("pluperfect", "sbjv"):["issem", "issēs", "isset", "issēmus", "issētis", "issent"],
}
def _idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _present_system(conj, pstem, tense, mood, voice, person, number):
"""Generate a present-system form (present/imperfect/future ind & subj)."""
table = _PARADIGM.get((tense, mood, voice))
if not table or conj not in table:
return None
return pstem + table[conj][_idx(person, number)]
def _active_infinitive_stem(conj, pstem):
return {1: pstem + "ā", 2: pstem + "ē", 3: pstem + "e",
"3io": pstem + "e", 4: pstem + "ī"}[conj]
_IRREG = {
"sum": {
("present", "ind", "active"): ["sum", "es", "est", "sumus", "estis", "sunt"],
("imperfect", "ind", "active"): ["eram", "erās", "erat", "erāmus", "erātis", "erant"],
("future", "ind", "active"): ["erō", "eris", "erit", "erimus", "eritis", "erunt"],
("perfect", "ind", "active"): ["fuī", "fuistī", "fuit", "fuimus", "fuistis", "fuērunt"],
("pluperfect", "ind", "active"): ["fueram", "fuerās", "fuerat", "fuerāmus", "fuerātis", "fuerant"],
("present", "sbjv", "active"): ["sim", "sīs", "sit", "sīmus", "sītis", "sint"],
("imperfect", "sbjv", "active"): ["essem", "essēs", "esset", "essēmus", "essētis", "essent"],
},
"possum": {
("present", "ind", "active"): ["possum", "potes", "potest", "possumus", "potestis", "possunt"],
("imperfect", "ind", "active"): ["poteram", "poterās", "poterat", "poterāmus", "poterātis", "poterant"],
("future", "ind", "active"): ["poterō", "poteris", "poterit", "poterimus", "poteritis", "poterunt"],
("perfect", "ind", "active"): ["potuī", "potuistī", "potuit", "potuimus", "potuistis", "potuērunt"],
("present", "sbjv", "active"): ["possim", "possīs", "possit", "possīmus", "possītis", "possint"],
},
"": {
("present", "ind", "active"): ["", "īs", "it", "īmus", "ītis", "eunt"],
("imperfect", "ind", "active"): ["ībam", "ībās", "ībat", "ībāmus", "ībātis", "ībant"],
("future", "ind", "active"): ["ībō", "ībis", "ībit", "ībimus", "ībitis", "ībunt"],
("perfect", "ind", "active"): ["", "īstī", "iit", "iimus", "īstis", "iērunt"],
("present", "sbjv", "active"): ["eam", "eās", "eat", "eāmus", "eātis", "eant"],
},
"volō": {
("present", "ind", "active"): ["volō", "vīs", "vult", "volumus", "vultis", "volunt"],
("imperfect", "ind", "active"): ["volēbam", "volēbās", "volēbat", "volēbāmus", "volēbātis", "volēbant"],
("future", "ind", "active"): ["volam", "volēs", "volet", "volēmus", "volētis", "volent"],
("perfect", "ind", "active"): ["voluī", "voluistī", "voluit", "voluimus", "voluistis", "voluērunt"],
("present", "sbjv", "active"): ["velim", "velīs", "velit", "velīmus", "velītis", "velint"],
},
"nōlō": {
("present", "ind", "active"): ["nōlō", "nōn vīs", "nōn vult", "nōlumus", "nōn vultis", "nōlunt"],
("present", "sbjv", "active"): ["nōlim", "nōlīs", "nōlit", "nōlīmus", "nōlītis", "nōlint"],
},
"ferō": {
("present", "ind", "active"): ["ferō", "fers", "fert", "ferimus", "fertis", "ferunt"],
("imperfect", "ind", "active"): ["ferēbam", "ferēbās", "ferēbat", "ferēbāmus", "ferēbātis", "ferēbant"],
("future", "ind", "active"): ["feram", "ferēs", "feret", "ferēmus", "ferētis", "ferent"],
("perfect", "ind", "active"): ["tulī", "tulistī", "tulit", "tulimus", "tulistis", "tulērunt"],
("present", "sbjv", "active"): ["feram", "ferās", "ferat", "ferāmus", "ferātis", "ferant"],
},
}
def conjugate(lemma, tense, mood, voice="active", person="third", number="singular"):
"""Return (surface, confidence). Perfect-passive forms are periphrastic and
handled in the realizer (sum + PPP); this returns synthetic forms only."""
lemma = lemma.strip()
i = _idx(person, number)
ir = _IRREG.get(lemma)
if ir:
tbl = ir.get((tense, mood, voice)) or ir.get((tense, mood, "active"))
if tbl and tbl[i]:
return tbl[i], "rule"
v = _VERBS.get(lemma)
if not v:
v = _infer_principal_parts(lemma)
if not v:
return lemma, "fallback"
conj, pstem, perfstem, supstem = v
# imperative (present active) 2sg / 2pl
if mood == "imp":
return _imperative(conj, pstem, person, number), "rule"
# perfect-system active
if tense in ("perfect", "pluperfect", "futureperfect") and voice == "active":
if not perfstem:
return lemma, "fallback"
end = _PERF_ACT.get((tense, mood))
if end:
return perfstem + end[i], "rule"
# present-system (active + passive)
if tense in ("present", "imperfect", "future"):
form = _present_system(conj, pstem, tense, mood, voice, person, number)
if form:
return form, "rule"
return lemma, "fallback"
def _imperative(conj, pstem, person, number):
if number == "singular":
return {1: pstem + "ā", 2: pstem + "ē", 3: pstem + "e",
"3io": pstem + "e", 4: pstem + "ī"}[conj]
return {1: pstem + "āte", 2: pstem + "ēte", 3: pstem + "ite",
"3io": pstem + "ite", 4: pstem + "īte"}[conj]
def _infer_principal_parts(lemma):
"""OOV fallback: infer conjugation + stems from the 1sg-present citation form.
Perfect/supine stems are guessed regularly (often wrong for 3rd conj) and the
resulting forms are still returned as 'rule' but the realizer down-weights."""
if lemma.endswith("ō"):
base = lemma[:-1]
# can't distinguish conj from 1sg alone reliably; default by ending vowel
if base.endswith("i"):
return ("3io", base[:-1], base[:-1] + "īv", base[:-1] + "īt")
return (3, base, base + "s", base + "t")
return None
# ── PUBLIC: participles ─────────────────────────────────────────────────────────
def participle(lemma, kind, case="nom", gender="m", number="singular"):
"""kind: 'prs' (present active, -ns/-ntis), 'pfv' (perfect passive, -tus),
'fut' (future active, -tūrus). Declined as an adjective via rule endings.
Returns (form, conf)."""
v = _VERBS.get(lemma)
if not v:
return lemma, "fallback"
conj, pstem, perfstem, supstem = v
if kind == "pfv":
if not supstem:
return lemma, "fallback"
base = supstem[:-1] if supstem.endswith("t") or supstem.endswith("s") else supstem
stem = supstem # supine stem already ends in t/s: amāt- -> amātus
return _decline_us_a_um(stem, case, gender, number), "rule"
if kind == "fut":
if not supstem:
return lemma, "fallback"
return _decline_us_a_um(supstem + "ūr", case, gender, number), "rule"
if kind == "prs":
# present active participle: stem + ns (nom), stem + nt- (oblique), 3rd-decl
pv = {1: "ā", 2: "ē", 3: "ē", "3io": "", 4: ""}[conj]
ntstem = pstem + pv + "nt"
return _decline_pres_ptcp(pstem + pv, case, gender, number), "rule"
return lemma, "fallback"
def _decline_us_a_um(stem, case, gender, number):
"""Decline a -us/-a/-um adjective/participle stem (2-1-2 declension)."""
C = _CASE_MAP.get(case, case.upper())
end = {
("NOM", "m", "singular"): "us", ("NOM", "f", "singular"): "a", ("NOM", "n", "singular"): "um",
("GEN", "m", "singular"): "ī", ("GEN", "f", "singular"): "ae", ("GEN", "n", "singular"): "ī",
("DAT", "m", "singular"): "ō", ("DAT", "f", "singular"): "ae", ("DAT", "n", "singular"): "ō",
("ACC", "m", "singular"): "um", ("ACC", "f", "singular"): "am", ("ACC", "n", "singular"): "um",
("ABL", "m", "singular"): "ō", ("ABL", "f", "singular"): "ā", ("ABL", "n", "singular"): "ō",
("VOC", "m", "singular"): "e", ("VOC", "f", "singular"): "a", ("VOC", "n", "singular"): "um",
("NOM", "m", "plural"): "ī", ("NOM", "f", "plural"): "ae", ("NOM", "n", "plural"): "a",
("GEN", "m", "plural"): "ōrum", ("GEN", "f", "plural"): "ārum", ("GEN", "n", "plural"): "ōrum",
("DAT", "m", "plural"): "īs", ("DAT", "f", "plural"): "īs", ("DAT", "n", "plural"): "īs",
("ACC", "m", "plural"): "ōs", ("ACC", "f", "plural"): "ās", ("ACC", "n", "plural"): "a",
("ABL", "m", "plural"): "īs", ("ABL", "f", "plural"): "īs", ("ABL", "n", "plural"): "īs",
("VOC", "m", "plural"): "ī", ("VOC", "f", "plural"): "ae", ("VOC", "n", "plural"): "a",
}.get((C, gender, number), "us")
return stem + end
def _decline_pres_ptcp(stem, case, gender, number):
"""Present active participle (amāns, amantis) — 3rd-declension, stem+ns/nt."""
C = _CASE_MAP.get(case, case.upper())
if C == "NOM" and number == "singular":
return stem + "ns"
if C == "VOC" and number == "singular":
return stem + "ns"
base = stem + "nt"
end = {
("GEN", "singular"): "is", ("DAT", "singular"): "ī",
("ACC", "singular"): "em" if gender != "n" else "",
("ABL", "singular"): "e",
("NOM", "plural"): "ēs" if gender != "n" else "ia",
("GEN", "plural"): "ium", ("DAT", "plural"): "ibus",
("ACC", "plural"): "ēs" if gender != "n" else "ia",
("ABL", "plural"): "ibus", ("VOC", "plural"): "ēs",
}.get((C, number), "is")
if C == "ACC" and number == "singular" and gender == "n":
return stem + "ns"
return base + end
def infinitive(lemma, tense="present", voice="active"):
lemma = lemma.strip()
if lemma == "sum":
return ("esse", "rule") if tense == "present" else ("fuisse", "rule")
v = _VERBS.get(lemma)
if not v:
return lemma, "fallback"
conj, pstem, perfstem, supstem = v
if tense == "present":
if voice == "active":
return _active_infinitive_stem(conj, pstem).rstrip() + \
("re" if conj != 3 and conj != "3io" else "re"), "rule"
# passive present infinitive
base = {1: pstem + "ā", 2: pstem + "ē", 4: pstem + "ī"}.get(conj)
if base:
return base + "", "rule"
return pstem + "ī", "rule" # 3rd: regī
if tense == "perfect" and voice == "active" and perfstem:
return perfstem + "isse", "rule"
return lemma, "fallback"
def lexicon_stats():
return {
"noun_adj_source": "UniMorph Latin (github.com/unimorph/lat, CC-BY-SA 3.0)",
"verb_source": "rule-based 4-conjugation engine over curated attested "
"principal parts (UniMorph verb list is a 947-lemma sample "
"MISSING all core verbs — amō/sum/videō absent)",
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
"curated_verb_lemmas": len(_VERBS) + len(_IRREG),
"gender_inference": "declension-based (nom+gen endings) + curated exceptions",
}
if __name__ == "__main__":
import json
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
print("\n-- noun declension puella (1st, fem) --")
for c in ("nom", "gen", "dat", "acc", "abl", "voc"):
print(f" {c}: sg={decline_noun('puella', c, 'singular')[0]:10} "
f"pl={decline_noun('puella', c, 'plural')[0]}")
print("\n-- rēx (3rd, m):", [decline_noun('rēx', c, 'singular')[0] for c in ('nom','gen','dat','acc','abl')])
print("-- gender: puella=", noun_gender("puella"), "rēx=", noun_gender("rēx"),
"bellum=", noun_gender("bellum"), "corpus=", noun_gender("corpus"),
"manus=", noun_gender("manus"), "diēs=", noun_gender("diēs"))
print("\n-- conjugate videō (2nd) present ind active --")
for p in ("first", "second", "third"):
for n in ("singular", "plural"):
print(f" {p[:3]}.{n[:2]}: {conjugate('videō','present','ind','active',p,n)[0]}")
print("-- amō forms:", conjugate("amō","present","ind","active","first","singular")[0],
conjugate("amō","imperfect","ind","active","third","plural")[0],
conjugate("amō","future","ind","active","first","singular")[0],
conjugate("amō","perfect","ind","active","third","singular")[0])
print("-- sum:", [conjugate("sum","present","ind","active",p,"singular")[0] for p in ("first","second","third")])
print("-- participle amō pfv acc.f.sg:", participle("amō","pfv","acc","f","singular")[0])
print("-- infinitive amō:", infinitive("amō")[0], "| regō pass:", infinitive("regō", voice="passive")[0])
-538
View File
@@ -1,538 +0,0 @@
"""morphology_pt_full.py — production-grade Brazilian-Portuguese morphological generator.
NOT a toy. Backed by two real, broad, Wiktionary-lineage lexicons:
VERBS UniMorph Portuguese (github.com/unimorph/por, CC-BY-SA 3.0)
4,001 verb lemmas × full paradigm (283,991 finite/non-finite forms +
20,005 participle forms). Every mood/tense pt actually inflects:
indicative present / preterite (PST;PFV) / imperfect (PST;IPFV) /
pluperfect-simple (PST;PRF) / future,
conditional (futuro do pretérito),
subjunctive present / imperfect / FUTURE (PT-specific live tense),
affirmative + negative imperative,
PERSONAL infinitive (V;{p};{n};NFIN a PT-specific finite-ish form),
past participle (4 gender/number forms) + gerúndio (V.PTCP;PRS).
NOUNS + ADJECTIVES kaikki.org Portuguese (Wiktionary extract, same lineage)
81,138 noun lemmas WITH inherent gender + real (often irregular) plural
so -ão-ões / -ãos / -ães / -õos is resolved PER LEMMA by Wiktionary,
never guessed (mãomãos, pãopães, coraçãocorações).
40,252 adjective lemmas with real feminine + masc/fem plural forms.
Fallbacks (degrade, never crash, on out-of-vocabulary input):
verbs : rule generator for regular -ar/-er/-ir paradigms
nouns : gender heuristic (endings) + rule pluralization (with -ão FLAGGED)
adjs : -o/-a gender rule + rule pluralization
Confidence flag on every form:
"lexicon" straight from UniMorph/kaikki (trust: high)
"rule" deterministic rule (trust: medium)
"fallback" could not inflect; returned lemma (trust: low -> FLAG)
Public API (used by realizer_pt.py):
conjugate(lemma, mood, tense, person, number) -> (form, conf)
personal_infinitive(lemma, person, number) -> (form, conf)
participle(lemma, gender="m", number="singular") -> (form, conf)
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"
inflect_noun(lemma, number, gender=None) -> (form, conf)
inflect_adj(lemma, gender, number) -> (form, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "por.unimorph")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_pt.jsonl")
_CACHE = os.path.join(_HERE, "data", "pt_morph_cache.pkl")
# ── mood/tense pair -> UniMorph feature triple (a in tag; b in tag; c in tag) ────
_VERB_KEYMAP = {
("ind", "present"): ("IND", "PRS", None),
("ind", "preterite"): ("IND", "PST", "PFV"),
("ind", "imperfect"): ("IND", "PST", "IPFV"),
("ind", "pluperfect"): ("IND", "PST", "PRF"), # simple mais-que-perfeito
("ind", "future"): ("IND", "FUT", None),
("ind", "conditional"): ("COND", None, None),
("sbjv", "present"): ("SBJV", "PRS", None),
("sbjv", "imperfect"): ("SBJV", "PST", "IPFV"),
("sbjv", "future"): ("SBJV", "FUT", None), # PT-specific
("imp", "affirmative"): ("IMP", "POS", None),
("imp", "negative"): ("IMP", "NEG", None),
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── build the compact lexicon from UniMorph (verbs) + kaikki (nouns/adjs) ────────
def _build_verbs():
verbs = {} # (lemma, "mood|tense|person|number") -> form
pinf = {} # (lemma, "person|number") -> personal-infinitive form
part = {} # lemma -> {("m","SG"): form, ...} past participle
ger = {} # lemma -> gerúndio
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f: # past participle: falado/falada/falados/faladas
g = "m" if "MASC" in f else ("f" if "FEM" in f else "m")
num = "SG" if "SG" in f else ("PL" if "PL" in f else "SG")
part.setdefault(lemma, {})[(g, num)] = form
elif "PRS" in f: # gerúndio: falando
ger.setdefault(lemma, form)
continue
if head != "V":
continue
# personal / impersonal infinitive
if "NFIN" in f:
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person and number:
pinf[(lemma, f"{person}|{number}")] = form
continue
# finite forms
mt = None
for (mood, tense), (a, b, c) in _VERB_KEYMAP.items():
if a not in f:
continue
if b is not None and b not in f:
continue
if c is not None and c not in f:
continue
# IND;PST needs exactly PFV|IPFV|PRF — reject if the required one absent
mt = (mood, tense)
break
if mt is None:
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
verbs.setdefault((lemma, f"{mt[0]}|{mt[1]}|{person}|{number}"), form)
return verbs, pinf, part, ger
def _kaikki_gender(arg):
if not arg:
return None
a = arg.lower()
if a.startswith("f"):
return "f"
if a.startswith("m"):
return "m"
return None
def _build_nouns_adjs():
nouns = {} # lemma -> {"g","SG","PL"}
adjs = {} # lemma -> {("m","SG"),("f","SG"),("m","PL"),("f","PL")}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
pos = d.get("pos")
word = d.get("word", "")
if not word or " " in word: # skip multiword entries
continue
forms = d.get("forms", []) or []
if pos == "noun":
ht = d.get("head_templates") or []
g = None
if ht:
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
if g is None:
tags = d.get("tags") or []
if "feminine" in tags:
g = "f"
elif "masculine" in tags:
g = "m"
pl = None
for x in forms:
t = x.get("tags") or []
if "plural" in t and "alternative" not in t and "obsolete" not in t:
pl = x.get("form")
break
# first entry wins; but a later entry with a plural fills a gap
if word not in nouns:
nouns[word] = {"g": g, "SG": word, "PL": pl}
else:
cur = nouns[word]
if cur.get("g") is None and g:
cur["g"] = g
if not cur.get("PL") and pl:
cur["PL"] = pl
elif pos == "adj":
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word)
for x in forms:
t = set(x.get("tags") or [])
fm = x.get("form")
if not fm or ("alternative" in t) or ("obsolete" in t):
continue
if "comparative" in t or "superlative" in t or \
"diminutive" in t or "augmentative" in t:
continue
if "feminine" in t and "plural" in t:
d0[("f", "PL")] = fm
elif "masculine" in t and "plural" in t:
d0[("m", "PL")] = fm
elif "feminine" in t:
d0[("f", "SG")] = fm
elif "plural" in t: # invariant-gender adj (feliz -> felizes)
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
return nouns, adjs
def _build_cache():
verbs, pinf, part, ger = _build_verbs()
nouns, adjs = _build_nouns_adjs()
data = {"verbs": verbs, "pinf": pinf, "part": part, "ger": ger,
"nouns": nouns, "adjs": adjs}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
newest_src = max(os.path.getmtime(_UNIMORPH),
os.path.getmtime(_KAIKKI) if os.path.exists(_KAIKKI) else 0)
if os.path.getmtime(_CACHE) >= newest_src:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PINF, _PART, _GER, _NOUNS, _ADJS = (
_LEX["verbs"], _LEX["pinf"], _LEX["part"], _LEX["ger"],
_LEX["nouns"], _LEX["adjs"])
# ── regular-ending rule fallback (deterministic, last resort) ────────────────────
def _vclass(lemma):
return lemma[-2:] if lemma[-2:] in ("ar", "er", "ir") else None
def _stem(lemma):
return lemma[:-2]
# endings indexed [1sg,2sg,3sg,1pl,2pl,3pl]
_REG = {
("ind", "present", "ar"): ["o", "as", "a", "amos", "ais", "am"],
("ind", "present", "er"): ["o", "es", "e", "emos", "eis", "em"],
("ind", "present", "ir"): ["o", "es", "e", "imos", "is", "em"],
("ind", "preterite", "ar"): ["ei", "aste", "ou", "amos", "astes", "aram"],
("ind", "preterite", "er"): ["i", "este", "eu", "emos", "estes", "eram"],
("ind", "preterite", "ir"): ["i", "iste", "iu", "imos", "istes", "iram"],
("ind", "imperfect", "ar"): ["ava", "avas", "ava", "ávamos", "áveis", "avam"],
("ind", "imperfect", "er"): ["ia", "ias", "ia", "íamos", "íeis", "iam"],
("ind", "imperfect", "ir"): ["ia", "ias", "ia", "íamos", "íeis", "iam"],
("sbjv", "present", "ar"): ["e", "es", "e", "emos", "eis", "em"],
("sbjv", "present", "er"): ["a", "as", "a", "amos", "ais", "am"],
("sbjv", "present", "ir"): ["a", "as", "a", "amos", "ais", "am"],
("sbjv", "imperfect", "ar"): ["asse", "asses", "asse", "ássemos", "ásseis", "assem"],
("sbjv", "imperfect", "er"): ["esse", "esses", "esse", "êssemos", "êsseis", "essem"],
("sbjv", "imperfect", "ir"): ["isse", "isses", "isse", "íssemos", "ísseis", "issem"],
("sbjv", "future", "ar"): ["ar", "ares", "ar", "armos", "ardes", "arem"],
("sbjv", "future", "er"): ["er", "eres", "er", "ermos", "erdes", "erem"],
("sbjv", "future", "ir"): ["ir", "ires", "ir", "irmos", "irdes", "irem"],
}
# future & conditional attach to the FULL infinitive
_FUT = ["ei", "ás", "á", "emos", "eis", "ão"]
_COND = ["ia", "ias", "ia", "íamos", "íeis", "iam"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
st, i = _stem(lemma), _slot_idx(person, number)
if mood == "ind" and tense == "future":
return lemma + _FUT[i]
if mood == "ind" and tense == "conditional":
return lemma + _COND[i]
if mood == "imp": # affirmative tú/vocês imperative ~ subjunctive present
table = _REG.get(("sbjv", "present", vc))
if table and tense == "negative":
return st + table[i]
# affirmative 2sg = 3sg present indicative; others = subjunctive
pres = _REG.get(("ind", "present", vc))
if person == "second" and number == "singular":
return st + pres[2]
return st + table[i] if table else None
table = _REG.get((mood, tense, vc))
if table:
return st + table[i]
return None
# verified corrections to UniMorph data errors (each audited individually, not
# guessed). The three 1PL-present entries are glued-allomorph errors surfaced by a
# full-lexicon scan for a non-final "mos" in V;1;PL;IND;PRS forms (the ONLY three).
_VERB_FIX = {
("estar", "ind", "imperfect", "third", "plural"): "estavam", # was "estávam"
("estar", "ind", "present", "first", "plural"): "estamos", # was "estamosestámos"
("haver", "ind", "present", "first", "plural"): "havemos", # was "havemoshemos"
("ir", "ind", "present", "first", "plural"): "vamos", # was "vamosimos"
}
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
"""Return (surface, confidence). mood in ind|sbjv|imp; tense per _VERB_KEYMAP."""
lemma = lemma.strip().lower()
fix = _VERB_FIX.get((lemma, mood, tense, person, number))
if fix:
return fix, "lexicon"
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
if form:
# pt-BR normalization: UniMorph `por` carries the EUROPEAN spelling of
# the -ar 1pl PRETERITE (-ámos). Brazilian PT drops the accent
# (falámos->falamos, chegámos->chegamos) — 3,334/4,001 verbs affected.
if (mood == "ind" and tense == "preterite" and person == "first"
and number == "plural" and form.endswith("ámos")):
form = form[:-4] + "amos"
return form, "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r:
return r, "rule"
return lemma, "fallback"
def personal_infinitive(lemma, person, number):
"""PT personal (inflected) infinitive: para falarmos, ao chegarem."""
lemma = lemma.strip().lower()
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _PINF.get((lemma, f"{p}|{n}"))
if form:
return form, "lexicon"
# rule: infinitive + personal endings (-, -es, -, -mos, -des, -em)
end = {("first", "singular"): "", ("second", "singular"): "es",
("third", "singular"): "", ("first", "plural"): "mos",
("second", "plural"): "des", ("third", "plural"): "em"}.get((person, number), "")
return lemma + end, "rule"
# ── PUBLIC: participle + gerund ───────────────────────────────────────────────────
def participle(lemma, gender="m", number="singular"):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
d = _PART.get(lemma)
if d:
form = d.get((g, num)) or d.get(("m", "SG"))
if form:
return form, "lexicon"
if lemma.endswith("ar"):
base = lemma[:-2] + "ad"
elif lemma[-2:] in ("er", "ir"):
base = lemma[:-2] + "id"
else:
return lemma, "fallback"
suf = {"m|SG": "o", "f|SG": "a", "m|PL": "os", "f|PL": "as"}[f"{g}|{num}"]
return base + suf, "rule"
def gerund(lemma):
lemma = lemma.strip().lower()
if lemma in _GER:
return _GER[lemma], "lexicon"
if lemma.endswith("ar"):
return lemma[:-2] + "ando", "rule"
if lemma.endswith("er"):
return lemma[:-2] + "endo", "rule"
if lemma.endswith("ir"):
return lemma[:-2] + "indo", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
_FEM_SUF = ("ção", "são", "ção", "dade", "tade", "agem", "igem", "ugem", "gem",
"ez", "eza", "ice", "ície", "tude", "ude", "âncbefore")
_FEM_SUF = ("ção", "são", "dade", "tade", "agem", "gem", "eza", "ez", "ice",
"tude", "ude", "ância", "ência", "ínia")
_MASC_SUF = ("ema", "oma", "ama", "grama", "eta", "ão") # Greek -ma etc. (mostly m)
def _gender_heuristic(noun):
for suf in _FEM_SUF:
if noun.endswith(suf):
return "f"
if noun.endswith(("ema", "oma", "ama")): # problema, idioma, programa
return "m"
if noun.endswith("a") or noun.endswith("ã"):
return "f"
if noun.endswith("o") or noun.endswith(("l", "r", "z", "m", "u", "i")):
return "m"
return "m"
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g"):
return d["g"]
return _gender_heuristic(lemma)
_INVARIANT_PL_SUF = ("s",) # paroxytones ending -s are invariant (o lápis / os lápis)
def _rule_plural(noun):
"""Deterministic PT pluralization. Returns (form, ok) where ok=False flags an
ambiguous -ão that should lower confidence (the lexicon normally resolves it)."""
if not noun:
return noun, True
if noun.endswith("ão"):
return noun[:-2] + "ões", False # majority rule, but AMBIGUOUS -> flag
if noun.endswith("m"):
return noun[:-1] + "ns", True # homem->homens, jardim->jardins
if noun.endswith("al"):
return noun[:-2] + "ais", True
if noun.endswith("el"):
return noun[:-2] + "éis", True
if noun.endswith("ol"):
return noun[:-2] + "óis", True
if noun.endswith("ul"):
return noun[:-2] + "uis", True
if noun.endswith("il"):
return noun[:-2] + "is", True # stressed (funil->funis); unstressed rarer
if noun.endswith(("r", "z")):
return noun + "es", True # flor->flores, luz->luzes
if noun.endswith("s"):
# paroxytone -s (lápis, ônibus) invariant; oxytone -s (país) -> -es
return noun, True
if noun.endswith(("a", "e", "i", "o", "u", "á", "é", "í", "ó", "ú", "ã")):
return noun + "s", True
return noun + "s", True
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if number == "singular":
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
if d and d.get("PL"):
return d["PL"], "lexicon"
form, ok = _rule_plural(lemma)
return form, ("rule" if ok else "fallback")
# ── PUBLIC: adjective agreement ──────────────────────────────────────────────────
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
d = _ADJS.get(lemma)
if d:
form = d.get((g, num))
if form:
return form, "lexicon"
# build a missing plural from this gender's singular
sg = d.get((g, "SG")) or d.get(("m", "SG")) or lemma
if num == "PL":
pl, ok = _rule_plural(sg)
return pl, ("rule" if ok else "fallback")
return sg, "lexicon"
# rule fallback: -o/-a gender, then pluralize
a = lemma
if g == "f":
if a.endswith("o"):
a = a[:-1] + "a"
elif a.endswith(("ês", "or")) and not a.endswith("ior"):
a = a + "a" # português->portuguesa, trabalhador->..a
if num == "PL":
a, ok = _rule_plural(a)
return a, ("rule" if ok else "fallback")
return a, "rule"
def lexicon_stats():
return {
"verb_source": "UniMorph Portuguese (github.com/unimorph/por)",
"noun_adj_source": "kaikki.org Portuguese (Wiktionary extract)",
"license": "CC-BY-SA (Wiktionary-derived)",
"verb_forms": len(_VERBS),
"verb_lemmas": len({k[0] for k in _VERBS}),
"personal_infinitive_forms": len(_PINF),
"participle_lemmas": len(_PART),
"gerund_lemmas": len(_GER),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("falar", "ind", "present", "first", "singular", "falo"),
("comer", "ind", "present", "third", "plural", "comem"),
("partir", "ind", "present", "first", "plural", "partimos"),
("ser", "ind", "present", "third", "singular", "é"),
("ir", "ind", "preterite", "first", "singular", "fui"),
("ter", "ind", "future", "first", "singular", "terei"),
("fazer", "sbjv", "present", "first", "singular", "faça"),
("dormir", "ind", "present", "first", "singular", "durmo"),
("dar", "ind", "preterite", "third", "singular", "deu"),
("poder", "ind", "conditional", "first", "singular", "poderia"),
("fazer", "sbjv", "future", "third", "singular", "fizer"),
("estar", "ind", "present", "third", "singular", "está"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
ok += got == exp
print(f" {flag}{lemma:8} {mood}/{tense} {per[:3]}.{num[:2]} -> {got:14} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" gender: casa=", noun_gender("casa"), "problema=", noun_gender("problema"),
"mão=", noun_gender("mão"), "coração=", noun_gender("coração"),
"flor=", noun_gender("flor"))
print(" plural: mão->", inflect_noun("mão", "plural"),
"| pão->", inflect_noun("pão", "plural"),
"| animal->", inflect_noun("animal", "plural"),
"| coração->", inflect_noun("coração", "plural"))
print(" adj: bonito/f/sg->", inflect_adj("bonito", "f", "singular"),
"| feliz/m/pl->", inflect_adj("feliz", "m", "plural"),
"| português/f/sg->", inflect_adj("português", "f", "singular"))
print(" part: fazer/m/sg->", participle("fazer"), "| ger falar->", gerund("falar"))
print(" pinf falar 1pl->", personal_infinitive("falar", "first", "plural"))
-609
View File
@@ -1,609 +0,0 @@
# -*- coding: utf-8 -*-
"""morphology_ro_full.py — production-grade Romanian morphological generator.
Romanian is the BIG typological delta of the Romance family. The verb engine and
the confidence/fallback contract TRANSFER from the Italian sibling; the NOMINAL
system is genuinely new: Romanian has a SUFFIXED definite article, a preserved
NOM/ACC vs GEN/DAT case distinction, a NEUTER gender (masc-agreeing in SG,
fem-agreeing in PL), and a VOCATIVE. Those are grounded in real per-lemma data,
not guessed.
Real, Wiktionary-lineage lexical sources:
VERBS UniMorph Romanian (github.com/unimorph/ron, CC-BY-SA 3.0)
~1216 verb lemmas × paradigm, CLEAN orthography:
indicativ prezent / imperfect (PST;IPFV) / perfectul simplu (PST;PFV) /
conjunctiv prezent (SBJV;PRS, stored WITHOUT the '' particle),
participiu (V.PTCP;PST, INVARIABLE in the perfect compus),
gerunziu (V.CVB;PRS), infinitiv (NFIN), imperativ.
ro_irreg_verbs (embedded) high-frequency verbs UniMorph MISSES
(avea, vrea, da) + the auxiliary clitic paradigms the compound tenses need
(perfect-compus am/ai/a/am/ați/au, viitor voi/vei/va/vom/veți/vor,
condițional /ai/ar/am/ați/ar). Real standard forms.
NOUNS kaikki.org Romanian (Wiktionary extract, CC-BY-SA 3.0)
the FULL declension per lemma, cleanly tagged:
(nom/acc | gen/dat | vocative) × (indefinite | definite) × (sg | pl).
This is what makes the suffixed article LEXICALLY grounded (omomul,
casăcasa, băiatbăiatul, casei gen/dat, omule vocative). Inherent gender
m / f / n (NEUTER available directly) from the head template.
ADJECTIVES UniMorph Romanian ADJ
full case × gender(MASC/FEM/NEUT) × number × definiteness paradigm.
Fallbacks (degrade, never crash, on OOV): rule verb conjugation for -a/-ea/-e/-i/-î
classes, rule pluralization, rule suffixed-article by gender+ending. Every form
carries a confidence flag: "lexicon" | "rule" | "fallback".
Public API (used by realizer_ro.py):
conjugate(lemma, mood, tense, person, number) -> (form, conf)
aux(kind, person, number) -> str # perfect / future / conditional clitics
participle(lemma) -> (form, conf) # INVARIABLE
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"|"n"
definite_suffix(noun, gender, number, case) -> (form, conf) # rule engine
inflect_noun(lemma, number, gender=None, case="nomacc", definite=False) -> (form, conf)
inflect_adj(lemma, gender, number, case="nomacc", definite=False) -> (form, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "ron.unimorph")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_ro.jsonl")
_CACHE = os.path.join(_HERE, "data", "ro_morph_cache.pkl")
# ── (mood, tense) -> UniMorph feature set ─────────────────────────────────────────
_VERB_KEYMAP = {
("ind", "present"): {"IND", "PRS"},
("ind", "imperfect"): {"IND", "PST", "IPFV"},
("ind", "perfect_s"): {"IND", "PST", "PFV"}, # perfectul simplu (regional/lit.)
("sbjv", "present"): {"SBJV", "PRS"},
("imp", "affirmative"): {"POS", "IMP"},
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── high-frequency irregulars UniMorph misses + auxiliary clitic paradigms ────────
# Real standard Romanian forms (textbook paradigms).
_IRREG = {
"avea": {
"ind|present|1|SG": "am", "ind|present|2|SG": "ai", "ind|present|3|SG": "are",
"ind|present|1|PL": "avem", "ind|present|2|PL": "aveți", "ind|present|3|PL": "au",
"ind|imperfect|1|SG": "aveam", "ind|imperfect|2|SG": "aveai",
"ind|imperfect|3|SG": "avea", "ind|imperfect|1|PL": "aveam",
"ind|imperfect|2|PL": "aveați", "ind|imperfect|3|PL": "aveau",
"sbjv|present|3|SG": "aibă", "sbjv|present|3|PL": "aibă",
"sbjv|present|1|SG": "am", "sbjv|present|2|SG": "ai",
"sbjv|present|1|PL": "avem", "sbjv|present|2|PL": "aveți",
"part": "avut", "ger": "având",
},
"vrea": {
"ind|present|1|SG": "vreau", "ind|present|2|SG": "vrei", "ind|present|3|SG": "vrea",
"ind|present|1|PL": "vrem", "ind|present|2|PL": "vreți", "ind|present|3|PL": "vor",
"ind|imperfect|1|SG": "voiam", "ind|imperfect|3|SG": "voia",
"sbjv|present|3|SG": "vrea", "sbjv|present|3|PL": "vrea",
"part": "vrut", "ger": "vrând",
},
"da": {
"ind|present|1|SG": "dau", "ind|present|2|SG": "dai", "ind|present|3|SG": "",
"ind|present|1|PL": "dăm", "ind|present|2|PL": "dați", "ind|present|3|PL": "dau",
"ind|imperfect|1|SG": "dădeam", "ind|imperfect|3|SG": "dădea",
"sbjv|present|3|SG": "dea", "sbjv|present|3|PL": "dea",
"part": "dat", "ger": "dând",
},
"fi": { # a fi — present is in UniMorph but keep participle + subjunctive here
"part": "fost", "ger": "fiind",
"sbjv|present|1|SG": "fiu", "sbjv|present|2|SG": "fii", "sbjv|present|3|SG": "fie",
"sbjv|present|1|PL": "fim", "sbjv|present|2|PL": "fiți", "sbjv|present|3|PL": "fie",
"ind|imperfect|1|SG": "eram", "ind|imperfect|2|SG": "erai",
"ind|imperfect|3|SG": "era", "ind|imperfect|1|PL": "eram",
"ind|imperfect|2|PL": "erați", "ind|imperfect|3|PL": "erau",
},
}
# auxiliary clitic paradigms (person,number)->form
_AUX = {
"perfect": {("first", "singular"): "am", ("second", "singular"): "ai",
("third", "singular"): "a", ("first", "plural"): "am",
("second", "plural"): "ați", ("third", "plural"): "au"},
"future": {("first", "singular"): "voi", ("second", "singular"): "vei",
("third", "singular"): "va", ("first", "plural"): "vom",
("second", "plural"): "veți", ("third", "plural"): "vor"},
"conditional": {("first", "singular"): "", ("second", "singular"): "ai",
("third", "singular"): "ar", ("first", "plural"): "am",
("second", "plural"): "ați", ("third", "plural"): "ar"},
}
def aux(kind, person, number):
return _AUX[kind][(person, number)]
# ── build verb lexicon from UniMorph ──────────────────────────────────────────────
def _build_verbs():
verbs, part, ger = {}, {}, {}
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f:
part.setdefault(lemma, form)
continue
if head == "V.CVB":
if "PRS" in f:
ger.setdefault(lemma, form)
continue
if head != "V":
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
# conjunctiv forms in UniMorph carry a leading 'să ' — strip it
surf = form
if surf.startswith(""):
surf = surf[3:]
for (mood, tense), req in _VERB_KEYMAP.items():
if not req <= f:
continue
if tense == "imperfect" and "PFV" in f:
continue
if tense == "perfect_s" and "IPFV" in f:
continue
# keep IND;PRS out of the PRF slot (mai-mult-ca-perfect etc. ignored)
if {"IND", "PRS"} <= req and "PRF" in f:
continue
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), surf)
break
return verbs, part, ger
# ── kaikki nouns: full declension paradigm per lemma ──────────────────────────────
_EXCL = {"alternative", "archaic", "obsolete", "regional", "dialectal", "rare",
"table-tags", "inflection-template", "error-unrecognized-form",
"diminutive", "augmentative", "informal"}
def _noun_key(tagset):
if tagset & _EXCL:
return None
if "vocative" in tagset:
case = "voc"
elif "genitive" in tagset or "dative" in tagset:
case = "gendat"
elif "nominative" in tagset or "accusative" in tagset:
case = "nomacc"
else:
return None
definite = "definite" in tagset and "indefinite" not in tagset
number = "PL" if "plural" in tagset else ("SG" if "singular" in tagset else None)
if number is None:
return None
return (case, definite, number)
def _build_nouns():
nouns = {} # lemma -> {"g":..., para:{(case,def,num):form}, "PL":plain_plural}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
if d.get("pos") != "noun":
continue
word = d.get("word", "")
if not word or " " in word:
continue
ht = d.get("head_templates") or []
g = None
if ht:
a = str((ht[0].get("args") or {}).get("1") or "").lower()
if a[:1] in ("m", "f", "n"):
g = a[:1]
entry = nouns.setdefault(word, {"g": g, "para": {}, "PL": None})
if entry["g"] is None and g:
entry["g"] = g
for x in (d.get("forms") or []):
fm = x.get("form")
tg = set(x.get("tags") or [])
if not fm or fm in ("-", "#", "") or " " in fm:
continue
if tg == {"plural"} and not entry["PL"]:
entry["PL"] = fm
k = _noun_key(tg)
if k and k not in entry["para"]:
entry["para"][k] = fm
return nouns
# ── adjectives from kaikki (UniMorph ron ADJ is sparse AND mis-tagged; kaikki is
# clean: the 4-form agreement pattern bun/bună/buni/bune). Neuter maps sg->masc,
# pl->fem, so 4 forms (m/f × SG/PL) fully cover it. ────────────────────────────
def _build_adjs():
adjs = {} # lemma -> {(gender,number): form} gender in {m,f}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
if d.get("pos") != "adj":
continue
word = d.get("word", "")
if not word or " " in word:
continue
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word) # masc sg = headword
for x in (d.get("forms") or []):
fm = x.get("form")
t = set(x.get("tags") or [])
if not fm or " " in fm or fm in ("-", "#") or (t & _EXCL):
continue
if "definite" in t or "genitive" in t or "dative" in t:
continue # keep indefinite nom/acc agr set
pl = "plural" in t
fem = "feminine" in t
masc = "masculine" in t
if fem and pl:
d0.setdefault(("f", "PL"), fm)
elif masc and pl:
d0.setdefault(("m", "PL"), fm)
elif fem and not pl:
d0.setdefault(("f", "SG"), fm)
elif pl and not fem and not masc: # bare plural -> both genders
d0.setdefault(("m", "PL"), fm)
d0.setdefault(("f", "PL"), fm)
return adjs
def _build_cache():
verbs, part, ger = _build_verbs()
nouns = _build_nouns()
adjs = _build_adjs()
data = {"verbs": verbs, "part": part, "ger": ger, "nouns": nouns, "adjs": adjs}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [_UNIMORPH, _KAIKKI]
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PART, _GER, _NOUNS, _ADJS = (
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"])
# ── rule verb conjugation fallback ────────────────────────────────────────────────
def _vclass(lemma):
if lemma.endswith("a"):
return "a"
if lemma.endswith("ea"):
return "ea"
if lemma.endswith("e"):
return "e"
if lemma.endswith("i"):
return "i"
if lemma.endswith("î"):
return "î"
return None
# regular present endings by class [1sg,2sg,3sg,1pl,2pl,3pl]
_REG_PRS = {
"a": ["", "i", "ă", "ăm", "ați", "ă"], # a lucra type (simplified)
"ea": ["", "i", "e", "em", "eți", "", ],
"e": ["", "i", "e", "em", "eți", ""],
"i": ["esc", "ești", "ește", "im", "iți", "esc"], # -i type (a vorbi)
"î": ["ăsc", "ăști", "ăște", "âm", "âți", "ăsc"],
}
_SLOT = {("first", "singular"): 0, ("second", "singular"): 1, ("third", "singular"): 2,
("first", "plural"): 3, ("second", "plural"): 4, ("third", "plural"): 5}
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
i = _SLOT[(person, number)]
body = lemma[:-len(vc)]
if mood == "ind" and tense == "present":
end = _REG_PRS[vc][i]
return body + end
if mood == "ind" and tense == "imperfect":
# -a/-i/-î -> stem + a/eai...; -e/-ea -> eam. Simplified regular imperfect.
stem = body
endings = {"a": ["am", "ai", "a", "am", "ați", "au"],
"i": ["eam", "eai", "ea", "eam", "eați", "eau"],
"î": ["am", "ai", "a", "am", "ați", "au"],
"e": ["eam", "eai", "ea", "eam", "eați", "eau"],
"ea": ["eam", "eai", "ea", "eam", "eați", "eau"]}[vc]
return stem + endings[i]
return None
# ── PUBLIC verb API ───────────────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
lemma = lemma.strip().lower()
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{_NUMBER.get(number,'?')}"
ir = _IRREG.get(lemma)
if ir and key in ir:
return ir[key], "lexicon"
form = _VERBS.get((lemma, key))
if form:
return form, "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r is not None:
return r, "rule"
return lemma, "fallback"
def participle(lemma):
"""Past participle — INVARIABLE in the perfect compus (am mers, am văzut)."""
lemma = lemma.strip().lower()
ir = _IRREG.get(lemma)
if ir and "part" in ir:
return ir["part"], "lexicon"
if lemma in _PART:
return _PART[lemma], "lexicon"
vc = _vclass(lemma)
if vc == "a":
return lemma[:-1] + "at", "rule"
if vc in ("ea",):
return lemma[:-2] + "ut", "rule"
if vc == "i":
return lemma[:-1] + "it", "rule"
if vc == "î":
return lemma[:-1] + "ât", "rule"
if vc == "e":
return lemma[:-1] + "ut", "rule"
return lemma, "fallback"
def gerund(lemma):
lemma = lemma.strip().lower()
ir = _IRREG.get(lemma)
if ir and "ger" in ir:
return ir["ger"], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
vc = _vclass(lemma)
if vc in ("a", "î"):
return lemma[:-1] + "ând", "rule"
if vc in ("ea", "e", "i"):
return lemma[:-len(vc)] + "ind", "rule"
return lemma, "fallback"
# ── noun gender ───────────────────────────────────────────────────────────────────
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g") in ("m", "f", "n"):
return d["g"]
if lemma.endswith(("ă", "a", "e")):
return "f"
return "m"
# ── SUFFIXED DEFINITE ARTICLE — rule engine (fallback for OOV nouns) ───────────────
def definite_suffix(noun, gender, number, case="nomacc"):
"""Attach the enclitic definite article by gender + ending. Returns (form, conf).
This is the headline Romanian-specific engine extension."""
n = noun
g = gender
if number == "singular":
if g in ("m", "n"):
if case == "gendat":
# masc/neut gen-dat definite: -lui
if n.endswith("e"):
return n + "lui", "rule" # câine -> câinelui
if n.endswith("u"):
return n + "lui", "rule"
return n + "ului", "rule" # om -> omului
# nom/acc
if n.endswith("e"):
return n + "le", "rule" # câine -> câinele
if n.endswith("u"):
return n + "l", "rule" # codru -> codrul
if n.endswith("i"):
return n + "ul", "rule"
return n + "ul", "rule" # om -> omul
# feminine singular
if case == "gendat":
# fem gen/dat definite = plural-stem + i (casei, fetei) — needs plural;
# approximated as: -ă->-ei, -e->-ei, -a->-alei
if n.endswith("ă"):
return n[:-1] + "ei", "rule" # casă -> casei
if n.endswith("e"):
return n[:-1] + "ei", "rule" # carte -> cărții(approx cartei)
if n.endswith("a"):
return n[:-1] + "lei", "rule"
return n + "i", "rule"
# fem nom/acc
if n.endswith("ă"):
return n[:-1] + "a", "rule" # casă -> casa
if n.endswith("e"):
return n[:-1] + "ea", "rule" # carte -> cartea
if n.endswith("a"):
return n + "ua", "rule" # stea -> steaua
if n.endswith("i"):
return n + "a", "rule"
return n + "a", "rule"
# plural
if case == "gendat":
base = noun
return base + "lor", "rule" # -lor for all gen/dat pl
if g == "m":
return noun + "i", "rule" # oameni -> oamenii (+i)
return noun + "le", "rule" # case -> casele, trenuri->trenurile
# ── rule pluralization (fallback) ─────────────────────────────────────────────────
def _rule_plural(noun, gender):
if gender == "f":
if noun.endswith("ă"):
return noun[:-1] + "e"
if noun.endswith("e"):
return noun[:-1] + "i"
if noun.endswith("a"):
return noun[:-1] + "le"
return noun + "e"
if gender == "n":
return noun + "uri"
# masculine
if noun.endswith(("e",)):
return noun[:-1] + "i"
return noun + "i"
# ── PUBLIC noun inflection ────────────────────────────────────────────────────────
def inflect_noun(lemma, number, gender=None, case="nomacc", definite=False):
lemma = lemma.strip().lower()
g = gender or noun_gender(lemma)
d = _NOUNS.get(lemma)
numk = "SG" if number == "singular" else "PL"
if d:
if case == "voc":
form = d["para"].get(("voc", True, numk)) or d["para"].get(("voc", False, numk))
if form:
return form, "lexicon"
# try the exact paradigm cell from kaikki (lexically grounded)
form = d["para"].get((case, definite, numk))
if form:
return form, "lexicon"
# indefinite fallbacks from the paradigm
if not definite:
form = d["para"].get(("nomacc", False, numk))
if form:
return form, "lexicon"
if numk == "PL" and d.get("PL"):
return d["PL"], "lexicon"
if numk == "SG":
return lemma, "lexicon"
# rule path
base = lemma if number == "singular" else _rule_plural(lemma, g)
if definite:
return definite_suffix(base, g, number, case)
return base, ("rule" if d is None else "lexicon")
# ── PUBLIC adjective agreement ────────────────────────────────────────────────────
def _neuter_map(gender, number):
# neuter agrees masculine in SG, feminine in PL
if gender == "n":
return "m" if number == "singular" else "f"
return gender
def inflect_adj(lemma, gender, number, case="nomacc", definite=False):
lemma = lemma.strip().lower()
numk = "SG" if number == "singular" else "PL"
eg = _neuter_map(gender, number) # neuter -> masc(SG)/fem(PL)
d = _ADJS.get(lemma)
if d:
form = d.get((eg, numk))
if form:
return form, "lexicon"
# rule fallback: 4-form pattern bun/bună/buni/bune keyed by effective gender
a = lemma
if number == "singular":
if eg == "f":
if a.endswith("e"):
return a, "rule" # mare invariant sg
if a.endswith("u"):
return a[:-1] + "ă", "rule" # nou -> nouă
if a.endswith("ă"):
return a, "rule"
return a + "ă", "rule" # bun -> bună
return a, "rule" # masc/neut sg = lemma
# plural
if eg == "f":
if a.endswith("e"):
return a[:-1] + "i", "rule" # mare -> mari
if a.endswith("u"):
return a[:-1] + "e", "rule" # nou -> noue (approx; 'noi' irr)
if a.endswith("ă"):
return a[:-1] + "e", "rule"
return a + "e", "rule" # bun -> bune
# masc/neut(SG-only)->here masc pl -> -i
if a.endswith("e"):
return a[:-1] + "i", "rule" # mare -> mari
if a.endswith("u"):
return a[:-1] + "i", "rule"
return a + "i", "rule" # bun -> buni
def lexicon_stats():
return {
"verb_source": "UniMorph Romanian (github.com/unimorph/ron) + curated "
"irregulars (avea/vrea/da + aux clitic paradigms)",
"noun_source": "kaikki.org Romanian — full case/definite/vocative declension",
"adj_source": "UniMorph Romanian ADJ (case×gender×number×definiteness)",
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
"unimorph_verb_forms": len(_VERBS),
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
"irregular_verb_lemmas": len(_IRREG),
"participle_lemmas": len(_PART),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
print("\n── SUFFIXED DEFINITE ARTICLE (the headline delta) ──")
for n, g in [("om", "m"), ("băiat", "m"), ("casă", "f"), ("carte", "f"),
("tren", "n"), ("student", "m"), ("floare", "f")]:
sg = inflect_noun(n, "singular", g, "nomacc", True)
pl = inflect_noun(n, "plural", g, "nomacc", True)
gd = inflect_noun(n, "singular", g, "gendat", True)
vo = inflect_noun(n, "singular", g, "voc", False)
print(f" {n:8}({g}) def.sg={sg[0]:12} def.pl={pl[0]:14} "
f"gen/dat.sg={gd[0]:12} voc={vo[0]}")
print("\n── NEUTER split agreement (tren: masc SG / fem PL) ──")
print(" tren nou ->", inflect_noun("tren", "singular", "n")[0],
inflect_adj("nou", "n", "singular")[0])
print(" trenuri noi->", inflect_noun("tren", "plural", "n")[0],
inflect_adj("nou", "n", "plural")[0])
print("\n── verbs ──")
for l, m, t, p, n, in [("merge", "ind", "present", "third", "singular"),
("avea", "ind", "present", "first", "singular"),
("fi", "ind", "present", "third", "singular"),
("vorbi", "ind", "present", "third", "plural"),
("face", "sbjv", "present", "third", "singular"),
("lucra", "ind", "imperfect", "third", "singular")]:
print(f" {l:8}{m}/{t:10}{p[:3]}.{n[:2]} -> {conjugate(l,m,t,p,n)}")
print(" perfect-aux(3sg):", aux("perfect", "third", "singular"),
"| future(1sg):", aux("future", "first", "singular"),
"| cond(3sg):", aux("conditional", "third", "singular"))
print(" participle merge/vedea:", participle("merge"), participle("vedea"))
-43
View File
@@ -1,43 +0,0 @@
// multilingual_gate.el - deterministic language detect + localized-phrase test.
fn mg_det(text: String, want: String) -> String {
let got: String = ml_detect(text)
let ok: String = "MISMATCH"
if str_eq(got, want) { let ok = "ok" }
return " detect(" + got + ") want=" + want + " (" + ok + ") :: " + text + "\n"
}
fn mg_ok(text: String, want: String) -> Int {
if str_eq(ml_detect(text), want) { return 1 }
return 0
}
fn run_ml_gate() -> String {
let t1: String = "Does Neuron use SQLite for storage?"
let t2: String = "Neuron, me explica cómo la saliencia forma las geometrías."
let t3: String = "O professor não leu o livro na memória."
let t4: String = "Che cosa memorizza Neuron nella memoria?"
let rep: String = "==== ELP multilingual detect + localized phrases ====\n"
let rep = rep + mg_det(t1, "en")
let rep = rep + mg_det(t2, "es")
let rep = rep + mg_det(t3, "pt")
let rep = rep + mg_det(t4, "it")
let rep = rep + " localized decline (pt): " + ml_tr("no_memory", "pt") + "\n"
let rep = rep + " localized decline (es): " + ml_tr("no_memory", "es") + "\n"
let rep = rep + " term(saliência->en): " + ml_term("saliência", "pt") + "\n"
let rep = rep + " pred(store->pt): " + ml_translate_pred("store", "pt") + "\n"
let ok: Int = 0
if mg_ok(t1, "en") == 1 { let ok = ok + 1 }
if mg_ok(t2, "es") == 1 { let ok = ok + 1 }
if mg_ok(t3, "pt") == 1 { let ok = ok + 1 }
if mg_ok(t4, "it") == 1 { let ok = ok + 1 }
let rep = rep + "-----------------------------------------------------------------\n"
let rep = rep + "language detected correctly: " + int_to_str(ok) + "/4\n"
if ok == 4 { let rep = rep + "ML GATE: PASS\n" } else { let rep = rep + "ML GATE: FAIL\n" }
return rep
}
println(run_ml_gate())
-52
View File
@@ -1,52 +0,0 @@
// propositions_gate.el - the READ primitive over memory text (native el).
// Proves triples are recovered from free memory text and that SACRED polarity
// survives extraction (a negative memory must yield a NOT-triple).
fn pg_check(text: String, want_pol: String) -> String {
let p: [String] = prop_extract_one(text, "nd-test")
let pol: String = slots_get(p, "polarity")
let ok: String = "MISMATCH"
if str_eq(pol, want_pol) { let ok = "ok" }
return " " + prop_repr(p) + " pol=" + pol + " expected=" + want_pol + " (" + ok + ")\n"
}
fn pg_pol_ok(text: String, want_pol: String) -> Int {
let p: [String] = prop_extract_one(text, "nd-test")
if str_eq(slots_get(p, "polarity"), want_pol) { return 1 }
return 0
}
fn run_prop_gate() -> String {
let m1: String = "Neuron stores memories in SQLite."
let m2: String = "The engram does not delete a memory."
let m3: String = "Salience never drops the negation."
let m4: String = "The teacher gives the book to the children."
let rep: String = "==== ELP proposition extraction (memory text -> triples) ====\n"
let rep = rep + pg_check(m1, "aff")
let rep = rep + pg_check(m2, "neg")
let rep = rep + pg_check(m3, "neg")
let rep = rep + pg_check(m4, "aff")
// multi-sentence memory: one triple per sentence, order preserved
let doc: String = "Neuron persists learning. It does not forget the library."
let props: [String] = prop_extract(doc, "nd-doc")
let rep = rep + " --- multi-sentence doc (" + int_to_str(native_list_len(props)) + " props) ---\n"
let di: Int = 0
while di < native_list_len(props) {
let rep = rep + " " + native_list_get(props, di) + "\n"
let di = di + 1
}
let ok: Int = 0
if pg_pol_ok(m1, "aff") == 1 { let ok = ok + 1 }
if pg_pol_ok(m2, "neg") == 1 { let ok = ok + 1 }
if pg_pol_ok(m3, "neg") == 1 { let ok = ok + 1 }
if pg_pol_ok(m4, "aff") == 1 { let ok = ok + 1 }
let rep = rep + "-----------------------------------------------------------------\n"
let rep = rep + "SACRED polarity correct on extraction: " + int_to_str(ok) + "/4\n"
if ok == 4 { let rep = rep + "PROP GATE: PASS\n" } else { let rep = rep + "PROP GATE: FAIL\n" }
return rep
}
println(run_prop_gate())
+13
View File
@@ -123,9 +123,22 @@ exec("EL_HTTP_TIMEOUT_MS=300000 " + SOME_BIN + " " + args + " 2>&1")
---
## Operator naming convention (for cognitive `.el` modules)
When you write `.el` that names a **cognitive faculty / operator** (the language
faculty, appreciation, interoception, the summon loop under `elp/`), name it for
its **functional human equivalent** — the faculty a mind would name — **not** its
linear-algebra operation. Put the math characterization in the `@impl`
doc-comment, never in the operator's public name (e.g. public `discern`/`contrast`
`@impl subtract (ab)`; public `summon`/`recall` ⟵ `@impl LOCAL nearest-region
+ bounded spreading activation`). Full table + rationale in
`foundation/el/AGENTS.md` § *Operator naming convention*. This does **not** apply
to plain library/compiler code (a `sort` is a `sort`).
## Rules
- New library functions → write in El
- New OS/hardware primitives → write in C and register in `codegen.el` arity table
- Never edit `dist/platform/elc` directly — always rebuild from source
- Never modify `el_seed.c` to add functionality that El can express
- Cognitive-faculty `.el` → name for the mind, algebra in `@impl` (see above)