Compare commits
34 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| ce34b94f88 | |||
| 0ae33c0f3b | |||
| c5508372ca | |||
| 335298a518 | |||
| 7d4fdbcc22 | |||
| 89ea1b5a15 | |||
| a816b119e7 | |||
| ba6e36c3f7 | |||
| 791b0880b7 | |||
| 23552ed40a | |||
| 6838e5cbff | |||
| fa2b49365b | |||
| 971b21751a | |||
| 9f1db8278c | |||
| 3d05e0c2a9 | |||
| 3bf44dee2d | |||
| a43a35bd10 | |||
| afc92f4e33 | |||
| 005e84e5d3 | |||
| 7f03876e26 | |||
| 599073cb92 | |||
| 7f66529510 | |||
| 6ebe3d0d66 | |||
| 9f362c90e5 | |||
| 11dc138a93 | |||
| 227f158a05 | |||
| 97e484221d | |||
| 8f8ccc945e | |||
| 409ec99397 | |||
| dc39a61e2c | |||
| eba9eac8a8 | |||
| ab6b52a0b4 | |||
| e3dabe3e08 | |||
| 0a0a2bcb44 |
@@ -0,0 +1,65 @@
|
||||
# ELP language consolidation — full-lexicon backfill (stage)
|
||||
|
||||
Branch: `stage-elp-lang-consolidation` (stage-bound; NOT the live soul :8742).
|
||||
|
||||
Consolidates scattered Python language-realizer work (`~/Desktop/lang-realizers`,
|
||||
`~/Desktop/lang-poetry-experiment`, `~/semitic_engine`) into the ELP `.el`
|
||||
structure, generating **full lexicons** (complete UniMorph + kaikki.org
|
||||
Wiktionary — real gender, real inflections) instead of the demo/curated subsets
|
||||
the prototypes shipped.
|
||||
|
||||
## ELP before this branch
|
||||
- 18 classical/ancient languages fully done (vocab + morphology + tests):
|
||||
akk ang cop egy enm fro gez goh got grc non peo pi sa sga sux txb uga.
|
||||
- 11 modern/classical languages had `morphology-<code>.el` in the build manifest
|
||||
but **no vocabulary and no lang_profile**: es fr de ja ar he hi ru fi sw la.
|
||||
- The ES port (`stage-elp-es-port`) had a *demo-scale* vocabulary-es.el (~350
|
||||
entries, s-expr form).
|
||||
|
||||
## Landed on this branch (full-lexicon seed-fn format, matching the 18 ancients)
|
||||
Vocabulary schema per row: `[lemma, pos, form0, form1, form2, en_gloss, hint]`.
|
||||
Files are ELP runtime **seed data** (loaded via the Engram at runtime), so — like
|
||||
all 18 classical `vocabulary-*.el` — they are intentionally NOT in the build
|
||||
manifest. Syntax validated: the chunked `fn vocab_<code>_seed_pN` format
|
||||
compiles cleanly to C via `elc` (correct UTF-8).
|
||||
|
||||
| code | in-ELP-morph? | vocab entries | verbs | nouns | adjs | profile |
|
||||
|------|---------------|--------------:|------:|------:|-----:|---------|
|
||||
| es | yes | 72,032 | 6,695 | 48,353 | 16,984 | yes |
|
||||
| fr | yes | 130,517 | 7,534 | 77,344 | 45,639 | yes |
|
||||
| de | yes | 144,692 | 6,661 | 133,162 | 4,869 | yes |
|
||||
| la | yes | 22,590 | 82 | 13,436 | 9,072 | yes |
|
||||
| it | no (bonus) | 193,675 | 10,008 | 109,459 | 74,208 | yes |
|
||||
| pt | no (bonus) | 115,772 | 4,001 | 72,073 | 39,698 | yes |
|
||||
| ro | no (bonus) | 86,504 | 1,216 | 65,915 | 19,373 | yes |
|
||||
| ca | no (bonus) | 47,112 | 1,547 | 28,830 | 16,735 | yes |
|
||||
|**total**| |**812,894** | | | | |
|
||||
|
||||
Generators (reproducible): `elp/tests/lang-gen/gen_elp_seed_full.py` (Romance),
|
||||
`gen_elp_seed_de_la.py` (German declension + Latin case-paradigm mapping). They
|
||||
read the pre-built morph caches in `~/Desktop/lang-realizers/data/` (UniMorph +
|
||||
kaikki), which are too large to commit.
|
||||
|
||||
## Remaining (honest)
|
||||
Of the 11 ELP backfill targets, 4 are done (es fr de la). The other 7 have **no
|
||||
full-lexicon engine** yet — cannot be generated honestly without engine work:
|
||||
- **ru**: only a 110-entry curated Slavic subset exists; full `rus.unimorph`
|
||||
present but no `morphology_ru_full` productive loader. Needs a full Russian
|
||||
morphology module (like the Romance ones) before vocab generation.
|
||||
- **ja / ko / zh**: validated demo engines (~66-104 hardcoded words) in
|
||||
`lang-poetry-experiment`, Python only. Agglutinative (ja/ko) + isolating (zh)
|
||||
need `.el` engine ports + full-lexicon wiring (ja: jpn_unimorph; zh: CC-CEDICT).
|
||||
- **ar / he (Semitic)**: template engines (16 AR / 8 HE patterns, ~6 roots) in
|
||||
`~/semitic_engine`, Python only. Root-and-pattern; full UniMorph ara/heb
|
||||
present but used only for validation. Needs productive root lexicon + `.el` port.
|
||||
- **hi (Hindi), fi (Finnish), sw (Swahili)**: `morphology-<code>.el` exists in
|
||||
ELP but there is NO scattered prototype and NO downloaded data for these —
|
||||
full-lexicon collection (UniMorph/kaikki) + generator still to do.
|
||||
|
||||
De/nl/sv Germanic and it/ro/ca/pt Romance verb coverage note: German verbs here
|
||||
are the ~6.6k caches carry; the it/ro/ca/pt bonus languages have full vocab but
|
||||
**no `morphology-<code>.el` in ELP yet** (Python realizer exists; `.el` port is
|
||||
the remaining engine work).
|
||||
|
||||
Construction coverage (separate from lexicon): French realizer was ~55%,
|
||||
Semitic ~3% in the prototypes — full construction coverage remains its own task.
|
||||
@@ -80,6 +80,11 @@ build {
|
||||
"src/grammar.el",
|
||||
"src/realizer.el",
|
||||
"src/semantics.el",
|
||||
"src/comprehend.el",
|
||||
"src/propositions.el",
|
||||
"src/multilingual.el",
|
||||
"src/self_region.el",
|
||||
"src/dialogue.el",
|
||||
"src/elp.el",
|
||||
]
|
||||
}
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,16 @@
|
||||
// comprehend.elh — public surface of the ELP comprehension front-end.
|
||||
// text → meaning-spec (the input half of the ELP; inverse of the realizer).
|
||||
extern fn parse_spec(text: String) -> [String]
|
||||
extern fn parse_spec_lang(text: String, lang: String) -> [String]
|
||||
extern fn parse_json(text: String) -> String
|
||||
extern fn parse_json_lang(text: String, lang: String) -> String
|
||||
// Analysis primitives (invertible morphology + deterministic grammar helpers):
|
||||
extern fn cp_tokenize(text: String) -> [String]
|
||||
extern fn cp_pron_concept(w: String) -> String
|
||||
extern fn cp_is_negation(w: String) -> Bool
|
||||
extern fn cp_is_neg_adverb(w: String) -> Bool
|
||||
extern fn cp_irr2(surface: String) -> [String]
|
||||
extern fn cp_reg_verb(w: String) -> [String]
|
||||
extern fn cp_analyze_verb(surface: String) -> [String]
|
||||
extern fn cp_verb_start(toks: [String], end: Int) -> Int
|
||||
extern fn cp_subord_start(toks: [String], n: Int) -> Int
|
||||
@@ -0,0 +1,287 @@
|
||||
// dialogue.el — SUMMON-THROUGH-SELF, native el. Port of dialogue.py's core.
|
||||
//
|
||||
// THE WHOLE DIALOGUE IS ONE OPERATION. A fact is never merely *fetched*: the
|
||||
// query is PROJECTED into the engram's self + memory geometry, LANDS in a region,
|
||||
// and the reply is READ OUT / the region MATERIALIZED from wherever it landed.
|
||||
//
|
||||
// project(query) -> land on a region -> read out from that region
|
||||
//
|
||||
// • lands in the SELF region -> grounded identity/presence, read out of
|
||||
// the real self nodes (self_region.el)
|
||||
// • lands on a memory NEIGHBORHOOD -> MATERIALIZE it: walk the neighborhood
|
||||
// (engram_neighbors_json) and read out the
|
||||
// region's connected members
|
||||
// • lands nowhere close -> HONEST ABSENCE (an empty region, not a
|
||||
// fabricated answer, not an error)
|
||||
//
|
||||
// CRITICAL INVARIANTS (enforced structurally, not by convention):
|
||||
// * ONE operation — there is NO intent classifier and NO separate
|
||||
// fact-retrieval branch. Identity is nearest-region proximity, not a switch.
|
||||
// * MATERIALIZE by walking the neighborhood, never by fetching top-props.
|
||||
// * HONEST ABSENCE when the region is thin.
|
||||
// * NEGATION is SACRED: the readout is the stored prose VERBATIM, so a negated
|
||||
// memory stays negated — we never paraphrase a polarity away.
|
||||
// * NO ECHO: the old "I noted that X. That relates to Y." template is gone.
|
||||
// The summon path materializes or honestly declines — it never echoes.
|
||||
// * DIRECTIVE OVERRIDE: a meta-directive ("answer in English") overrides the
|
||||
// reply language while the content language is still auto-detected.
|
||||
//
|
||||
// Depends on: comprehend (parse_spec_lang, cp_tokenize), multilingual (ml_detect,
|
||||
// ml_tr, ml_term), propositions (prop_split_sentences), self_region
|
||||
// (sr_available, sr_readout), the engram + json runtime builtins.
|
||||
|
||||
// ── directive override ────────────────────────────────────────────────────────
|
||||
// Return [target_lang, content]. target_lang is "" when no directive is present.
|
||||
// A directive names an output language; we strip it and keep the remaining text
|
||||
// as the content (whose OWN language is still auto-detected downstream).
|
||||
|
||||
fn dlg_dir_hit(low: String, phrase: String) -> Bool {
|
||||
return str_contains(low, phrase)
|
||||
}
|
||||
|
||||
fn dlg_parse_directive(text: String) -> [String] {
|
||||
let low: String = str_to_lower(text)
|
||||
let lang: String = ""
|
||||
let phrase: String = ""
|
||||
// English target
|
||||
if dlg_dir_hit(low, "in english") { let lang = "en"; let phrase = "in english" }
|
||||
if dlg_dir_hit(low, "em inglês") { let lang = "en"; let phrase = "em inglês" }
|
||||
if dlg_dir_hit(low, "em ingles") { let lang = "en"; let phrase = "em ingles" }
|
||||
if dlg_dir_hit(low, "en inglés") { let lang = "en"; let phrase = "en inglés" }
|
||||
// Portuguese target
|
||||
if dlg_dir_hit(low, "in portuguese") { let lang = "pt"; let phrase = "in portuguese" }
|
||||
if dlg_dir_hit(low, "em português") { let lang = "pt"; let phrase = "em português" }
|
||||
// Spanish target
|
||||
if dlg_dir_hit(low, "in spanish") { let lang = "es"; let phrase = "in spanish" }
|
||||
if dlg_dir_hit(low, "en español") { let lang = "es"; let phrase = "en español" }
|
||||
// Italian target
|
||||
if dlg_dir_hit(low, "in italian") { let lang = "it"; let phrase = "in italian" }
|
||||
|
||||
let content: String = text
|
||||
if !str_eq(phrase, "") {
|
||||
// strip the directive phrase (and a common "answer"/"responda" lead-in),
|
||||
// leaving the real question as content.
|
||||
let idx: Int = str_index_of(low, phrase)
|
||||
if idx >= 0 {
|
||||
let before: String = str_slice(text, 0, idx)
|
||||
let after: String = str_slice(text, idx + str_len(phrase), str_len(text))
|
||||
let content = str_trim(before + " " + after)
|
||||
}
|
||||
// trim a leading "answer"/"responda"/"reply" and stray colon/comma.
|
||||
let cl: String = str_to_lower(content)
|
||||
if str_starts_with(cl, "answer") { let content = str_trim(str_slice(content, 6, str_len(content))) }
|
||||
if str_starts_with(cl, "responda") { let content = str_trim(str_slice(content, 8, str_len(content))) }
|
||||
if str_starts_with(cl, "reply") { let content = str_trim(str_slice(content, 5, str_len(content))) }
|
||||
if str_starts_with(content, ":") { let content = str_trim(str_slice(content, 1, str_len(content))) }
|
||||
if str_starts_with(content, ",") { let content = str_trim(str_slice(content, 1, str_len(content))) }
|
||||
}
|
||||
let r: [String] = native_list_empty()
|
||||
let r = native_list_append(r, lang)
|
||||
let r = native_list_append(r, content)
|
||||
return r
|
||||
}
|
||||
|
||||
// ── identity landing (a region proximity, not a classifier switch) ────────────
|
||||
// The query lands in the SELF region when it takes an identity/presence shape.
|
||||
// Cross-lingual forms are included because the engram's lexical probe is
|
||||
// English-leaning. This is the SELF attractor of the single operation.
|
||||
|
||||
fn dlg_is_identity(content: String) -> Bool {
|
||||
let low: String = str_to_lower(str_trim(content))
|
||||
if str_contains(low, "who are you") { return true }
|
||||
if str_contains(low, "what are you") { return true }
|
||||
if str_contains(low, "who i am") { return true }
|
||||
if str_contains(low, "your name") { return true }
|
||||
if str_contains(low, "about yourself") { return true }
|
||||
if str_contains(low, "are you conscious") { return true }
|
||||
if str_contains(low, "are you there") { return true }
|
||||
// cross-lingual identity question-forms
|
||||
if str_contains(low, "quem é você") { return true }
|
||||
if str_contains(low, "quem es voce") { return true }
|
||||
if str_contains(low, "quién eres") { return true }
|
||||
if str_contains(low, "quien eres") { return true }
|
||||
if str_contains(low, "chi sei") { return true }
|
||||
if str_contains(low, "qui es-tu") { return true }
|
||||
if str_contains(low, "wer bist du") { return true }
|
||||
return false
|
||||
}
|
||||
|
||||
// ── readout helpers ───────────────────────────────────────────────────────────
|
||||
|
||||
fn dlg_first_sentence(content: String) -> String {
|
||||
let sents: [String] = prop_split_sentences(content)
|
||||
let n: Int = native_list_len(sents)
|
||||
let i: Int = 0
|
||||
while i < n {
|
||||
let s: String = str_trim(native_list_get(sents, i))
|
||||
// drop a leading markdown heading marker for a clean read-out line
|
||||
if str_starts_with(s, "# ") { let s = str_trim(str_slice(s, 2, str_len(s))) }
|
||||
if str_len(s) > 0 { return s }
|
||||
let i = i + 1
|
||||
}
|
||||
return str_trim(content)
|
||||
}
|
||||
|
||||
// strip trailing/leading punctuation from a token.
|
||||
fn dlg_clean_tok(w: String) -> String {
|
||||
let s: String = str_trim(w)
|
||||
let s = str_strip_suffix(s, ".")
|
||||
let s = str_strip_suffix(s, ",")
|
||||
let s = str_strip_suffix(s, "?")
|
||||
let s = str_strip_suffix(s, "!")
|
||||
let s = str_strip_suffix(s, ":")
|
||||
let s = str_strip_suffix(s, ";")
|
||||
return str_trim(s)
|
||||
}
|
||||
|
||||
// closed-class across the supported languages (union) — a word we must NOT treat
|
||||
// as a retrieval topic. Also drops the meta verbs of a request ("tell", "prove",
|
||||
// "show") so the TOPIC, not the speech act, is what projects into memory.
|
||||
fn dlg_is_stop(w: String) -> Bool {
|
||||
if ml_stop_en(w) { return true }
|
||||
if ml_stop_es(w) { return true }
|
||||
if ml_stop_pt(w) { return true }
|
||||
if ml_stop_it(w) { return true }
|
||||
if str_eq(w, "tell") { return true }
|
||||
if str_eq(w, "show") { return true }
|
||||
if str_eq(w, "about") { return true }
|
||||
if str_eq(w, "sobre") { return true }
|
||||
if str_eq(w, "acerca") { return true }
|
||||
return false
|
||||
}
|
||||
|
||||
// The CONTENT TERMS the query projects into memory: content words only, cleaned,
|
||||
// cross-lingually mapped to the engram's English vocabulary, ≥3 chars. This is
|
||||
// the geometry probe — the speech-act verbs and function words are stripped so a
|
||||
// PP topic ("tell me ABOUT Lisbon") projects on "lisbon", not "tell"/"me".
|
||||
fn dlg_content_terms(content: String, lang: String) -> [String] {
|
||||
let toks: [String] = cp_tokenize(content)
|
||||
let n: Int = native_list_len(toks)
|
||||
let out: [String] = native_list_empty()
|
||||
let i: Int = 0
|
||||
while i < n {
|
||||
let w: String = str_to_lower(dlg_clean_tok(native_list_get(toks, i)))
|
||||
if str_len(w) >= 3 {
|
||||
if !dlg_is_stop(w) {
|
||||
let out = native_list_append(out, ml_term(w, lang))
|
||||
}
|
||||
}
|
||||
let i = i + 1
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// Does this landed node lexically overlap the query's content terms? This is the
|
||||
// RELEVANCE FLOOR: activation always returns the store's most salient nodes, so
|
||||
// without this a query about nothing would "land" on the self/top node. A node
|
||||
// that shares no content term with the query is "nowhere close" -> honest absence.
|
||||
fn dlg_node_matches(node: String, terms: [String]) -> Bool {
|
||||
let hay: String = str_to_lower(json_get_string(node, "content") + " " + json_get_string(node, "label"))
|
||||
let n: Int = native_list_len(terms)
|
||||
let i: Int = 0
|
||||
while i < n {
|
||||
let t: String = native_list_get(terms, i)
|
||||
if str_len(t) >= 3 {
|
||||
if str_contains(hay, t) { return true }
|
||||
}
|
||||
let i = i + 1
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// MATERIALIZE the landed region: read out the landed fact, then WALK the
|
||||
// neighborhood and read out its connected members (real edges, not top-props).
|
||||
fn dlg_materialize(top_node: String, reply_lang: String) -> String {
|
||||
let id: String = json_get_string(top_node, "id")
|
||||
let content: String = json_get_string(top_node, "content")
|
||||
let lead: String = dlg_first_sentence(content)
|
||||
|
||||
let nb: String = engram_neighbors_json(id, 2, "both")
|
||||
let m: Int = json_array_len(nb)
|
||||
let parts: [String] = native_list_empty()
|
||||
let parts = native_list_append(parts, lead)
|
||||
let added: Int = 0
|
||||
let i: Int = 0
|
||||
while i < m {
|
||||
if added < 3 {
|
||||
let rec: String = json_array_get(nb, i)
|
||||
let node: String = json_get_raw(rec, "node")
|
||||
let nc: String = json_get_string(node, "content")
|
||||
if !str_eq(nc, "") {
|
||||
let sent: String = dlg_first_sentence(nc)
|
||||
if !str_eq(sent, "") {
|
||||
let parts = native_list_append(parts, sent)
|
||||
let added = added + 1
|
||||
}
|
||||
}
|
||||
}
|
||||
let i = i + 1
|
||||
}
|
||||
// The readout is the region's OWN prose, verbatim — negation SACRED, no echo.
|
||||
return str_join(parts, " ")
|
||||
}
|
||||
|
||||
// ── THE single operation ──────────────────────────────────────────────────────
|
||||
|
||||
fn dlg_respond(text: String) -> String {
|
||||
// directive override: reply language may differ from content language.
|
||||
let dir: [String] = dlg_parse_directive(text)
|
||||
let target_lang: String = native_list_get(dir, 0)
|
||||
let content: String = native_list_get(dir, 1)
|
||||
|
||||
let content_lang: String = ml_detect(content)
|
||||
let reply_lang: String = content_lang
|
||||
if !str_eq(target_lang, "") { let reply_lang = target_lang }
|
||||
|
||||
// comprehend the content (SACRED polarity carried in the spec).
|
||||
let spec: [String] = parse_spec_lang(content, content_lang)
|
||||
|
||||
// ── PROJECT + LAND: SELF region ───────────────────────────────────────────
|
||||
// Identity/presence shape lands in the self region; read out the REAL self
|
||||
// nodes (self_region.el), never a template. Same single operation — this is
|
||||
// just the self attractor winning the landing.
|
||||
if dlg_is_identity(content) {
|
||||
if sr_available() {
|
||||
// read out the REAL self nodes when replying in their own language
|
||||
// (the soul's prose is English); for another reply language we cannot
|
||||
// translate real content without an LLM, so we answer with the
|
||||
// localized SACRED identity anchor — honest, in-language, no fabrication.
|
||||
if str_eq(reply_lang, "en") { return sr_readout("en") }
|
||||
return ml_tr("identity", reply_lang)
|
||||
}
|
||||
// self region thin — honest localized identity (logged fallback shape).
|
||||
return ml_tr("identity", reply_lang)
|
||||
}
|
||||
|
||||
// ── PROJECT into MEMORY geometry ──────────────────────────────────────────
|
||||
let terms: [String] = dlg_content_terms(content, content_lang)
|
||||
let qterm: String = str_join(terms, " ")
|
||||
let act: String = engram_activate_json(qterm, 12)
|
||||
let n: Int = json_array_len(act)
|
||||
|
||||
// ── LAND: the highest-activation node that ACTUALLY overlaps the query's
|
||||
// content terms (the relevance floor). Activation always returns the most
|
||||
// salient nodes, so we walk the ranked list and take the first that is
|
||||
// genuinely "close"; if none is, the query landed nowhere. ───────────────
|
||||
let landing: String = ""
|
||||
let i: Int = 0
|
||||
while i < n {
|
||||
if str_eq(landing, "") {
|
||||
let rec: String = json_array_get(act, i)
|
||||
let node: String = json_get_raw(rec, "node")
|
||||
if dlg_node_matches(node, terms) {
|
||||
let landing = node
|
||||
}
|
||||
}
|
||||
let i = i + 1
|
||||
}
|
||||
|
||||
// ── HONEST ABSENCE: nothing close — an empty region, not a fabricated answer,
|
||||
// not an "I noted that" echo. ────────────────────────────────────────────
|
||||
if str_eq(landing, "") {
|
||||
return ml_tr("no_memory", reply_lang)
|
||||
}
|
||||
|
||||
// ── MATERIALIZE the landing by WALKING its neighborhood. ──────────────────
|
||||
return dlg_materialize(landing, reply_lang)
|
||||
}
|
||||
@@ -63,6 +63,9 @@ import "morphology-cop.el"
|
||||
import "grammar.el"
|
||||
import "realizer.el"
|
||||
import "semantics.el"
|
||||
|
||||
// ── Comprehension front-end (input half: text → meaning-spec) ─────────────────
|
||||
import "comprehend.el"
|
||||
//
|
||||
// Entry points:
|
||||
//
|
||||
@@ -117,6 +120,9 @@ fn build_form_from_json(semantic_form_json: String, lang_code: String) -> [Strin
|
||||
let location: String = sem_get(semantic_form_json, "location")
|
||||
let tense: String = sem_get(semantic_form_json, "tense")
|
||||
let aspect: String = sem_get(semantic_form_json, "aspect")
|
||||
let polarity: String = sem_get(semantic_form_json, "polarity")
|
||||
let neg_word: String = sem_get(semantic_form_json, "neg_word")
|
||||
let iobj: String = sem_get(semantic_form_json, "iobj")
|
||||
|
||||
let form: [String] = native_list_empty()
|
||||
let form = native_list_append(form, "intent")
|
||||
@@ -127,12 +133,19 @@ fn build_form_from_json(semantic_form_json: String, lang_code: String) -> [Strin
|
||||
let form = native_list_append(form, predicate)
|
||||
let form = native_list_append(form, "patient")
|
||||
let form = native_list_append(form, patient)
|
||||
let form = native_list_append(form, "iobj")
|
||||
let form = native_list_append(form, iobj)
|
||||
let form = native_list_append(form, "location")
|
||||
let form = native_list_append(form, location)
|
||||
let form = native_list_append(form, "tense")
|
||||
let form = native_list_append(form, tense)
|
||||
let form = native_list_append(form, "aspect")
|
||||
let form = native_list_append(form, aspect)
|
||||
// SACRED: polarity crosses the JSON boundary and is never inferred away.
|
||||
let form = native_list_append(form, "polarity")
|
||||
let form = native_list_append(form, polarity)
|
||||
let form = native_list_append(form, "neg_word")
|
||||
let form = native_list_append(form, neg_word)
|
||||
let form = native_list_append(form, "lang")
|
||||
let form = native_list_append(form, lang_code)
|
||||
|
||||
|
||||
@@ -0,0 +1,72 @@
|
||||
;;; lang_profile_ca.el — Catalan language profile for ELP.
|
||||
;;; Mirrors lang_profile_it / _es / _pt; keys the realizer's construction switches.
|
||||
;;; Catalan is the CLOSEST Romance sibling to the shared engine (~85% conceptual
|
||||
;;; reuse). The deltas: PRONOMS FEBLES with four position allomorphs, l'-elision,
|
||||
;;; del/al/pel contractions, the periphrastic preterite (vaig+INF), and NO
|
||||
;;; essere/avere split (perfect aux is always HAVER; ser/estar is only the copula).
|
||||
|
||||
(lang_profile_ca
|
||||
(language "Catalan")
|
||||
(iso639 "ca")
|
||||
(family "Romance")
|
||||
|
||||
;; ── core typology flags ────────────────────────────────────────────────
|
||||
(pro-drop yes) ; null subjects default; overt pronoun = emphatic
|
||||
(obligatory-subject no)
|
||||
(grammatical-gender yes) ; m/f; full NP agreement (art + adj + participle)
|
||||
(do-support no)
|
||||
(subject-aux-inversion no) ; yes/no Q = declarative order + '?'; no inversion
|
||||
(article-selection "el/la/l'/els/les ; un/una/uns/unes") ; l'-ELISION:
|
||||
; el/la -> l' before vowel or (silent) h, glued to
|
||||
; the next word (l'home, l'illa); de -> d' before vowel
|
||||
(article-drives-contraction yes) ; article choice feeds prep+article contraction
|
||||
(adjective-position "postnominal-default + small prenominal class") ; bo/bon,
|
||||
; mal, gran, nou, vell, primer, molt... prenominal
|
||||
(question-punct plain) ; ? and ! only (no inverted ¿ ¡)
|
||||
|
||||
;; ── MANDATORY prep+article contractions ────────────────────────────────
|
||||
(contractions ((de el del) (de els dels)
|
||||
(a el al) (a els als)
|
||||
(per el pel) (per els pels)))
|
||||
(contraction-mandatory yes) ; *de el -> del obligatory
|
||||
(contraction-blocked-before-elision yes) ; de l'home / a l'home (NO *del home)
|
||||
|
||||
;; ── clitic system: PRONOMS FEBLES (the headline delta) ──────────────────
|
||||
(clitics yes)
|
||||
(clitic-allomorphy four-position) ; per pronoun, form varies by position+onset:
|
||||
; reinforced (em, et, el) proclitic before a consonant
|
||||
; elided (m', t', l', n') proclitic before a vowel/h
|
||||
; full (-me, -lo, -li) enclitic after a consonant/-r
|
||||
; reduced ('m, 't, 'l, 'ns) enclitic after a vowel
|
||||
(clitic-placement ((finite proclitic) ; el veig, no m'ho dóna
|
||||
(imperative-affirmative enclitic) ; dóna'm, digues-me
|
||||
(imperative-negative present-subjunctive) ; no parlis (delta)
|
||||
(infinitive enclitic) ; ajudar-me, veure'l
|
||||
(gerund enclitic))) ; fent-ho
|
||||
(clitic-combination ((me el "me'l") (te el "te'l") (se el "se'l")
|
||||
(me la "me la") (me en "me'n")
|
||||
(li el "l'hi") (li en "n'hi"))) ; dative+accusative clusters
|
||||
(clitic-particles (hi en ho)) ; locative hi, partitive/genitive en, neuter ho
|
||||
|
||||
;; ── verb / aspect system ───────────────────────────────────────────────
|
||||
(finite-agreement "person+number (6-way)")
|
||||
(tenses (present imperfet preterit-simple perifrastic-preterit futur
|
||||
condicional subjuntiu-present subjuntiu-imperfet imperatiu))
|
||||
(periphrastic-preterite "vaig/vas/va/vam/vau/van + INFINITIVE") ; << hallmark CA
|
||||
; (vaig cantar = 'I sang'); coexists w/ synthetic pret.
|
||||
(compound-past "pretèrit perfet = haver(present) + participle")
|
||||
(perfect-aux "HAVER only") ; << NO essere/avere split (simpler than IT)
|
||||
(participle-agreement ((haver preceding-acc-clitic))) ; les he vistes; else invariable
|
||||
(progressive-aux "estar + gerundi")
|
||||
(copula "ser / estar") ; ser: identity/essential/origin; estar:
|
||||
; location + transient state (estic cansat, és a casa)
|
||||
(passive-aux "ser (+ per-agent)")
|
||||
(future inflectional) ; cantaré, serà
|
||||
(comparative "més/menys ADJ que")
|
||||
|
||||
;; ── SACRED safety bar (shared with es/pt/it/en) ────────────────────────
|
||||
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
|
||||
(negation "no (preverbal) + optional 'pas' + concord") ; no...res/
|
||||
; ningú/mai/cap/gens/enlloc
|
||||
(negative-concord yes) ; preverbal negative subject (ningú) keeps 'no'
|
||||
(neg-reinforcer pas)) ; optional (no ho faré pas)
|
||||
@@ -0,0 +1,41 @@
|
||||
;;; lang_profile_de.el — German language profile for ELP.
|
||||
;;; Mirrors lang_profile_en / lang_profile_es. Keys the realizer's construction
|
||||
;;; switches. German is the largest Germanic delta from the EN engine: V2 word
|
||||
;;; order, four morphological cases, and separable-prefix verbs.
|
||||
|
||||
(lang_profile_de
|
||||
(language "German")
|
||||
(iso639 "de")
|
||||
(family "Germanic")
|
||||
(neighbor-base "en") ; realized by extending the English (Germanic) engine
|
||||
|
||||
;; ── core typology flags ────────────────────────────────────────────────
|
||||
(pro-drop no) ; obligatory subject in finite clauses
|
||||
(obligatory-subject yes)
|
||||
(grammatical-gender (m f n)) ; three genders; drives article + adj declension
|
||||
(case-system (nom acc dat gen)) ; four cases on articles/adjs/nouns
|
||||
(word-order V2) ; finite verb 2nd in main clause
|
||||
(subordinate-order verb-final) ; "..., dass er den Hund SIEHT."
|
||||
(separable-verbs yes) ; aufstehen -> "steht ... auf"; ppart "aufgestanden"
|
||||
(do-support no) ; German negates/questions the finite verb directly
|
||||
(subject-verb-inversion yes) ; yes/no Q fronts finite verb; wh-Q fills Vorfeld
|
||||
(article-selection "der/die/das + ein/kein") ; declined by case x gender x number
|
||||
(adjective-position prenominal)
|
||||
(adjective-declension (strong weak mixed)) ; chosen by the determiner type
|
||||
(noun-capitalization yes)
|
||||
|
||||
;; ── verb / aspect system ───────────────────────────────────────────────
|
||||
(finite-agreement "person-and-number") ; full present/past paradigm
|
||||
(auxiliary-order (modal tense-aux perfect passive main))
|
||||
(perfect-aux (haben sein)) ; sein for intransitive motion/change verbs
|
||||
(passive-aux "werden")
|
||||
(future "werden + infinitive")
|
||||
(comparative "synthetic (-er / -st, with umlaut)")
|
||||
|
||||
;; ── negation ───────────────────────────────────────────────────────────
|
||||
(negation-markers (nicht kein)) ; kein- negates an indefinite NP; nicht else
|
||||
(negation-faithful yes) ; SACRED: polarity never dropped/inverted -> FLAG
|
||||
|
||||
;; ── lexicon provenance ─────────────────────────────────────────────────
|
||||
(lexicon-source "UniMorph deu (primary) + kaikki.org German (gender override)")
|
||||
(lexicon-license "CC-BY-SA 3.0 / GFDL"))
|
||||
@@ -0,0 +1,41 @@
|
||||
;;; lang_profile_en.el — English language profile for ELP.
|
||||
;;; Mirrors lang_profile_es / lang_profile_pt; keys the realizer's construction
|
||||
;;; switches. English is typologically distinct from the Romance builds, so the
|
||||
;;; flags differ where the grammar differs.
|
||||
|
||||
(lang_profile_en
|
||||
(language "English")
|
||||
(iso639 "en")
|
||||
(family "Germanic")
|
||||
|
||||
;; ── core typology flags ────────────────────────────────────────────────
|
||||
(pro-drop no) ; OBLIGATORY subjects — missing subject is FLAGGED
|
||||
(obligatory-subject yes)
|
||||
(grammatical-gender no) ; natural gender only (he/she/it), no NP agreement
|
||||
(do-support yes) ; negation & questions of lexical verbs insert do/does/did
|
||||
(subject-aux-inversion yes) ; yes/no + non-subject wh questions invert the operator
|
||||
(article-selection "a/an/the") ; a/an resolved PHONOLOGICALLY (an hour, a university)
|
||||
(adjective-position prenominal) ; attributive adjectives precede the noun; invariant
|
||||
(has-tag-questions yes) ; "...doesn't he?" — operator + reversed polarity
|
||||
(has-there-existential yes) ; "there is/are/have been ..."
|
||||
(possessive-clitic "'s") ; saxon genitive; plural in -s -> bare apostrophe
|
||||
(question-punct plain) ; ? and ! only (no inverted marks)
|
||||
|
||||
;; ── verb / aspect system ───────────────────────────────────────────────
|
||||
(finite-agreement "3sg-present-only") ; only 3sg present -s (+ suppletive be)
|
||||
(auxiliary-order (modal perfect progressive passive main))
|
||||
(perfect-aux "have") ; have + past participle
|
||||
(progressive-aux "be") ; be + present participle
|
||||
(passive-aux "be") ; be + past participle (+ by-agent)
|
||||
(future "will + base") ; no inflectional future
|
||||
(comparative "synthetic-or-periphrastic") ; -er/-est vs more/most by syllables
|
||||
|
||||
;; ── SACRED safety bar (shared with es/pt) ──────────────────────────────
|
||||
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
|
||||
|
||||
;; ── DIALECT overlay (post-realization, one core -> US/UK/AU) ────────────
|
||||
(dialect US) ; default; profile field switches the overlay
|
||||
(dialects (US UK AU))
|
||||
(dialect-canonical US) ; core is authored in US orthography
|
||||
(dialect-overlay "dialect_en.to_dialect") ; orthography + lexis + grammar prefs
|
||||
(dialect-covers (spelling lexis collective-agreement gotten/got)))
|
||||
@@ -0,0 +1,45 @@
|
||||
;;; lang_profile_es.el — Spanish language profile for ELP.
|
||||
;;; Keys the realizer's construction switches. Mirrors lang_profile_en / _pt.
|
||||
|
||||
(lang_profile_es
|
||||
(language "Spanish")
|
||||
(iso639 "es")
|
||||
(family "Romance")
|
||||
|
||||
;; -- core typology flags -------------------------------------------------
|
||||
(pro-drop yes) ; subjects routinely dropped; agreement carries person
|
||||
(obligatory-subject no)
|
||||
(grammatical-gender yes) ; m/f on every noun; article+adjective AGREE
|
||||
(gender-source lexicon); REAL per-noun gender from UniMorph — NOT a heuristic
|
||||
(do-support no)
|
||||
(subject-aux-inversion no) ; questions by intonation/punctuation, not inversion
|
||||
(question-strategy intonation)
|
||||
(article-selection "el/la/los/las un/una/unos/unas")
|
||||
(stressed-a-rule yes) ; fem sg noun in stressed a-/ha- takes el/un (el agua)
|
||||
(adjective-position postnominal) ; default post; a few prenominal + apocope
|
||||
(adjective-agreement "gender+number")
|
||||
(question-punct inverted) ; opening ¿ ¡ required
|
||||
|
||||
;; -- MANDATORY CONTRACTIONS (coordinator quality bar) --------------------
|
||||
(contractions ((de el "del") (a el "al")))
|
||||
(contraction-mandatory yes) ; 'de el'/'a el' MUST surface as del/al
|
||||
|
||||
;; -- verb / aspect system ------------------------------------------------
|
||||
(verb-classes (ar er ir))
|
||||
(tenses (present preterite imperfect future conditional))
|
||||
(moods (ind sbjv imp))
|
||||
(finite-agreement "person+number (6 slots)")
|
||||
(perfect-aux "haber") ; haber + past participle (invariant -o)
|
||||
(progressive-aux "estar") ; estar + gerund
|
||||
(passive-aux "ser") ; ser + participle (agrees) + por-agent
|
||||
(copula-split "ser/estar") ; permanent vs stage-level
|
||||
(future "infinitive + é/ás/á/emos/éis/án")
|
||||
|
||||
;; -- clitics / government ------------------------------------------------
|
||||
(object-clitics yes) ; me te lo la le nos os los las; proclisis/enclisis
|
||||
(clitic-order "se II I III (le+lo -> se lo)")
|
||||
(enclisis "imperative/infinitive/gerund + accent repair (dá+me+lo->dámelo)")
|
||||
(verb-prep-government yes) ; verbs select prep (protestar+contra, escapar+de)
|
||||
|
||||
;; -- SACRED safety bar (shared with en/pt) -------------------------------
|
||||
(negation-faithful yes)) ; polarity never dropped/inverted; unplaceable -> FLAG
|
||||
@@ -0,0 +1,74 @@
|
||||
;;; lang_profile_fr.el — French language profile for ELP.
|
||||
;;; Mirrors lang_profile_it / lang_profile_es; keys the realizer's construction
|
||||
;;; switches. French is a Romance sibling (~54% of the realizer code and the whole
|
||||
;;; clause-engine architecture reused), but carries the family's biggest surface
|
||||
;;; deltas: NOT pro-drop, DISCONTINUOUS negation, and an orthography/phonology
|
||||
;;; mismatch (elision, liaison) that makes exact-match genuinely hard.
|
||||
|
||||
(lang_profile_fr
|
||||
(language "French")
|
||||
(iso639 "fr")
|
||||
(family "Romance")
|
||||
|
||||
;; ── core typology flags ────────────────────────────────────────────────
|
||||
(pro-drop no) ; << French-specific: subject clitic OBLIGATORY
|
||||
(obligatory-subject yes) ; je/tu/il/elle/nous/vous/ils/elles always overt
|
||||
(grammatical-gender yes) ; m/f; full NP agreement (art + adj + participle)
|
||||
(do-support no)
|
||||
(subject-aux-inversion optional) ; est-ce que (default) OR clitic inversion (vas-tu)
|
||||
(article-selection "le/la/l'/les ; un/une/des ; PARTITIVE du/de la/de l'/des")
|
||||
(article-drives-contraction yes) ; à+le=au, de+le=du feed off article choice
|
||||
(adjective-position "postnominal-default + prenominal-BAGS") ; beau/bon/grand/
|
||||
; petit/jeune/vieux/nouveau + ordinals prenominal
|
||||
; (beau->bel, nouveau->nouvel, vieux->vieil / vowel)
|
||||
(question-punct "space-before") ; French typography: ' ?' ' !' (no ¿¡)
|
||||
|
||||
;; ── elision (orthography/phonology mismatch — French-specific) ──────────
|
||||
(elision ((le l') (la l') (je j') (ne n') (de d') (que qu')
|
||||
(me m') (te t') (se s') (ce c'))) ; before vowel / h-muet
|
||||
(elision-h-muet yes) ; l'homme, l'hôpital (h-aspiré exception list kept)
|
||||
(liaison noted-not-modeled) ; phonological, not written in surface
|
||||
|
||||
;; ── MANDATORY prep+article contractions ────────────────────────────────
|
||||
(contractions ((à le au) (à les aux) (de le du) (de les des)))
|
||||
(contraction-mandatory yes) ; *à le -> au obligatory; à la / à l' uncontracted
|
||||
(partitive ((m-sg du) (f-sg "de la") (vowel "de l'") (pl des)))
|
||||
(partitive-under-neg "de") ; << gap in current build: 'ne … pas de pain'
|
||||
|
||||
;; ── clitic system ──────────────────────────────────────────────────────
|
||||
(clitics yes)
|
||||
(clitic-order (me te se nous vous | le la les | lui leur | y | en))
|
||||
(clitic-placement ((finite proclitic) ; je le lui donne
|
||||
(imperative-affirmative enclitic-hyphen) ; donne-le-moi
|
||||
(imperative-negative "ne+proclitic+verb+pas") ; ne le donne pas
|
||||
(infinitive enclitic))) ; PARTIAL: clitic-climbing
|
||||
; onto infinitive under modal
|
||||
(clitic-imperative-shift ((me moi) (te toi))) ; final me/te -> moi/toi (donne-moi)
|
||||
(clitic-particles (y en)) ; locative y, partitive/genitive en
|
||||
|
||||
;; ── verb / aspect system ───────────────────────────────────────────────
|
||||
(finite-agreement "person+number (written; many homophones)")
|
||||
(tenses (présent imparfait passé-simple futur conditionnel
|
||||
subjonctif-présent subjonctif-imparfait impératif))
|
||||
(compound-past "passé-composé = aux(present) + participe passé")
|
||||
(perfect-aux "être/avoir (LEXICAL selection)") ; << French-specific
|
||||
(etre-aux-class "intransitive motion/change (aller venir arriver partir
|
||||
entrer sortir monter descendre naître mourir rester
|
||||
tomber retourner passer devenir revenir rentrer) + ALL
|
||||
pronominal verbs")
|
||||
(participle-agreement ((être subject) ; elle est allée / elles venues
|
||||
(avoir preceding-direct-object))) ; je les ai vus
|
||||
(progressive "être en train de + infinitif") ; no dedicated aux
|
||||
(copula "être (single; no ser/estar, no essere/stare)")
|
||||
(passive-aux "être (+ par-agent)")
|
||||
(future inflectional) ; parlera, sera
|
||||
(comparative "plus/moins ADJ que")
|
||||
(superlative "le/la plus ADJ (de …)") ; PARTIAL word-order in build
|
||||
|
||||
;; ── SACRED safety bar (shared with es/pt/it/en) ────────────────────────
|
||||
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
|
||||
(negation "DISCONTINUOUS: ne (preverbal) … pas/jamais/rien/personne/
|
||||
plus/guère/que (postverbal)") ; << biggest structural delta
|
||||
(negation-ne-elides yes) ; ne -> n' before vowel (n'ai pas vu)
|
||||
(negation-passe-composé "ne + aux + pas + participe") ; n'ai pas vu
|
||||
(negative-concord partial)) ; personne/rien as arguments post-participle
|
||||
@@ -0,0 +1,70 @@
|
||||
;;; lang_profile_it.el — Italian language profile for ELP.
|
||||
;;; Mirrors lang_profile_es / lang_profile_pt; keys the realizer's construction
|
||||
;;; switches. Italian is a Romance sibling, so ~85% of the flags match ES/PT; the
|
||||
;;; essere/avere auxiliary split and phonological article selection are the deltas.
|
||||
|
||||
(lang_profile_it
|
||||
(language "Italian")
|
||||
(iso639 "it")
|
||||
(family "Romance")
|
||||
|
||||
;; ── core typology flags ────────────────────────────────────────────────
|
||||
(pro-drop yes) ; null subjects default; overt pronoun = emphatic
|
||||
(obligatory-subject no)
|
||||
(grammatical-gender yes) ; m/f; full NP agreement (art + adj + participle)
|
||||
(do-support no)
|
||||
(subject-aux-inversion no) ; yes/no Q = declarative order + '?'; no inversion
|
||||
(article-selection "il/lo/l'/i/gli + la/l'/le ; un/uno/un'/una") ; PHONOLOGICAL:
|
||||
; lo/gli/uno before s+cons, z, gn, ps, pn, x, y, i+V;
|
||||
; l'/un' before a vowel (elision, glued to next word)
|
||||
(article-drives-contraction yes) ; article choice feeds the prep+art contraction
|
||||
(adjective-position "postnominal-default + prenominal-class") ; bello/buono/grande
|
||||
; /nuovo/vecchio/primo... prenominal (with apocope)
|
||||
(question-punct plain) ; ? and ! only (no inverted ¿ ¡)
|
||||
|
||||
;; ── MANDATORY prep+article contractions ────────────────────────────────
|
||||
(contractions ((di il del) (di lo dello) (di la della) (di i dei)
|
||||
(di gli degli) (di le delle) (di l' dell')
|
||||
(a il al) (a lo allo) (a la alla) (a i ai) (a gli agli)
|
||||
(a le alle) (a l' all')
|
||||
(da il dal) (da la dalla) (da gli dagli) (da l' dall')
|
||||
(in il nel) (in la nella) (in gli negli) (in l' nell')
|
||||
(su il sul) (su la sulla) (su gli sugli) (su l' sull')))
|
||||
(contraction-mandatory yes) ; *di il -> del is obligatory, never uncontracted
|
||||
(prep-no-contract (per tra fra)) ; per la strada (NOT *perla)
|
||||
|
||||
;; ── clitic system ──────────────────────────────────────────────────────
|
||||
(clitics yes)
|
||||
(clitic-placement ((finite proclitic) ; lo vedo, non me lo dà
|
||||
(imperative-affirmative enclitic) ; dammelo, guardalo
|
||||
(imperative-negative-tu non+infinitive) ; non parlare / non lo fare
|
||||
(infinitive enclitic) ; vederlo, aiutarmi (drop -e)
|
||||
(gerund enclitic))) ; dandolo
|
||||
(clitic-combination ((mi lo "me lo") (ti lo "te lo") (ci lo "ce lo")
|
||||
(vi lo "ve lo") (si lo "se lo")
|
||||
(gli lo "glielo") (le lo "glielo"))) ; glielo = ONE word
|
||||
(clitic-particles (ci ne)) ; locative ci, partitive ne
|
||||
(raddoppiamento (da fa di va sta)) ; monosyllabic imper double clitic: dammelo
|
||||
|
||||
;; ── verb / aspect system ───────────────────────────────────────────────
|
||||
(finite-agreement "person+number (6-way)")
|
||||
(tenses (presente imperfetto passato-remoto futuro condizionale
|
||||
congiuntivo-presente congiuntivo-imperfetto imperativo))
|
||||
(compound-past "passato-prossimo = aux(present) + participle")
|
||||
(perfect-aux "essere/avere (LEXICAL selection)") ; << Italian-specific
|
||||
(essere-aux-class unaccusative) ; motion/change-of-state/copular/pronominal
|
||||
; (andare venire nascere morire diventare piacere
|
||||
; + ALL reflexives) -> essere
|
||||
(participle-agreement ((essere subject) ; è andata / sono arrivati
|
||||
(avere preceding-acc-clitic))) ; li ho visti
|
||||
(progressive-aux "stare + gerundio") ; sto parlando
|
||||
(copula "essere (default) / stare (state: sto bene)")
|
||||
(passive-aux "essere / venire (+ da-agent)")
|
||||
(future inflectional) ; parlerò, sarà
|
||||
(comparative "più/meno ADJ di")
|
||||
|
||||
;; ── SACRED safety bar (shared with es/pt/en) ───────────────────────────
|
||||
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
|
||||
(negation "non (preverbal) + concord") ; non...niente/nessuno/mai/più
|
||||
(negative-concord yes) ; preverbal negative word (nessuno/niente) suppresses non
|
||||
(neg-adverb-position between-aux-and-participle)) ; non ho MAI visto
|
||||
@@ -0,0 +1,30 @@
|
||||
;;; lang_profile_la.el — Latin language profile for ELP.
|
||||
;;; Keys the realizer's construction switches. Companion to morphology-la.el.
|
||||
|
||||
(lang_profile_la
|
||||
(language "Latin")
|
||||
(iso639 "la")
|
||||
(family "Italic")
|
||||
|
||||
;; -- core typology flags -------------------------------------------------
|
||||
(pro-drop yes) ; person carried by verb ending; subjects dropped
|
||||
(obligatory-subject no)
|
||||
(grammatical-gender yes) ; m/f/n; adjective AGREES in case+gender+number
|
||||
(gender-source lexicon) ; REAL per-noun gender from UniMorph lat
|
||||
(articles none) ; Latin has no articles
|
||||
(case-system yes) ; NOM GEN DAT ACC ABL VOC (+ rare LOC)
|
||||
(cases (nom gen dat acc abl voc))
|
||||
(word-order "SOV (default; free order, case-marked)")
|
||||
(adjective-position "either (case agreement carries the link)")
|
||||
(adjective-agreement "case+gender+number")
|
||||
|
||||
;; -- verb / aspect system ------------------------------------------------
|
||||
(verb-classes (1 2 3 3io 4)) ; four conjugations + i-stem 3rd
|
||||
(tenses (present imperfect future perfect pluperfect futureperfect))
|
||||
(moods (indicative subjunctive imperative infinitive))
|
||||
(voices (active passive))
|
||||
(finite-agreement "person+number (6 slots)")
|
||||
(citation "principal parts: pres-1sg / pres-inf / perf-participle")
|
||||
|
||||
;; -- SACRED safety bar ---------------------------------------------------
|
||||
(negation-faithful yes)) ; polarity never dropped/inverted
|
||||
@@ -0,0 +1,40 @@
|
||||
;;; lang_profile_pt.el — Portuguese language profile for ELP.
|
||||
;;; Keys the realizer's construction switches. Mirrors lang_profile_es.
|
||||
|
||||
(lang_profile_pt
|
||||
(language "Portuguese")
|
||||
(iso639 "pt")
|
||||
(family "Romance")
|
||||
|
||||
;; -- core typology flags -------------------------------------------------
|
||||
(pro-drop yes) ; subjects routinely dropped; agreement carries person
|
||||
(obligatory-subject no)
|
||||
(grammatical-gender yes) ; m/f on every noun; article+adjective AGREE
|
||||
(gender-source lexicon) ; REAL per-noun gender from UniMorph por / kaikki
|
||||
(do-support no)
|
||||
(subject-aux-inversion no)
|
||||
(question-strategy intonation)
|
||||
(article-selection "o/a/os/as um/uma/uns/umas")
|
||||
(adjective-position postnominal)
|
||||
(adjective-agreement "gender+number")
|
||||
|
||||
;; -- MANDATORY CONTRACTIONS (prep + article) -----------------------------
|
||||
(contractions ((de o "do") (de a "da") (em o "no") (em a "na")
|
||||
(a o "ao") (a a "à") (por o "pelo") (por a "pela")))
|
||||
(contraction-mandatory yes)
|
||||
|
||||
;; -- verb / aspect system ------------------------------------------------
|
||||
(verb-classes (ar er ir))
|
||||
(tenses (present preterite imperfect future conditional))
|
||||
(moods (ind sbjv imp))
|
||||
(finite-agreement "person+number (6 slots)")
|
||||
(perfect-aux "ter") ; ter + past participle
|
||||
(copula-split "ser/estar")
|
||||
(personal-infinitive yes) ; distinctive PT inflected infinitive
|
||||
|
||||
;; -- clitics / government ------------------------------------------------
|
||||
(object-clitics yes) ; mesoclisis/enclisis/proclisis by context
|
||||
(verb-prep-government yes)
|
||||
|
||||
;; -- SACRED safety bar ---------------------------------------------------
|
||||
(negation-faithful yes))
|
||||
@@ -0,0 +1,71 @@
|
||||
;;; lang_profile_ro.el — Romanian language profile for ELP.
|
||||
;;; Romanian is the BIG typological delta of the Romance family. The verb/clause
|
||||
;;; engine and the SACRED negation contract mirror the ES/PT/IT core, but the
|
||||
;;; NOMINAL system is genuinely new: a SUFFIXED definite article, preserved CASE,
|
||||
;;; a NEUTER gender, and a VOCATIVE. Those flags mark where the shared engine was
|
||||
;;; extended rather than reused.
|
||||
|
||||
(lang_profile_ro
|
||||
(language "Romanian")
|
||||
(iso639 "ro")
|
||||
(family "Romance (Eastern / Balkan)")
|
||||
|
||||
;; ── core typology flags ────────────────────────────────────────────────
|
||||
(pro-drop yes) ; null subjects default; overt pronoun = emphatic
|
||||
(obligatory-subject no)
|
||||
(grammatical-gender yes) ; m / f / NEUTER (n)
|
||||
(neuter-gender yes) ; << ROMANIAN-SPECIFIC: masc-agreeing SG, fem-agreeing PL
|
||||
; (un tren nou / două trenuri noi)
|
||||
(do-support no)
|
||||
(subject-aux-inversion no) ; yes/no Q = declarative order + '?'
|
||||
(question-punct plain) ; ? and ! only
|
||||
|
||||
;; ── SUFFIXED DEFINITE ARTICLE (the headline engine extension) ───────────
|
||||
(definite-article suffixed) ; << UNIQUE IN ROMANCE: enclitic on the noun
|
||||
(definite-forms ((m/n sg "-ul / -le / -l : om->omul, câine->câinele, codru->codrul")
|
||||
(f sg "-a / -ea / -ua : casă->casa, carte->cartea, stea->steaua")
|
||||
(m pl "-i : oameni->oamenii")
|
||||
(f/n pl "-le : case->casele, trenuri->trenurile")))
|
||||
(article-host ((no-prenom-adj noun) ; omul bun
|
||||
(prenom-adj adjective))) ; bunul om (adj carries the article)
|
||||
(indefinite-article ((m/n "un") (f "o") (pl "niște") (gen/dat-pl "unor")))
|
||||
|
||||
;; ── CASE (preserved; NOM/ACC vs GEN/DAT) ────────────────────────────────
|
||||
(case (nom/acc gen/dat vocative)) ; << ROMANIAN-SPECIFIC
|
||||
(case-syncretism "nom=acc ; gen=dat")
|
||||
(genitive-marking "gen/dat definite: -lui (m/n), -ei/-i (f), -lor (pl)")
|
||||
(genitival-article ((m sg "al") (f sg "a") (m pl "ai") (f/n pl "ale"))) ; o carte a lui
|
||||
(possession "definite-head + gen/dat possessor: casa băiatului")
|
||||
(vocative ((m sg "-ule/-e : omule, băiete") (f sg "-o : Mario, fato")
|
||||
(pl "-lor")))
|
||||
|
||||
;; ── verb / aspect system ────────────────────────────────────────────────
|
||||
(finite-agreement "person+number (6-way)")
|
||||
(tenses (prezent imperfect perfect-simplu conjunctiv-prezent
|
||||
imperativ (periphrastic: perfect-compus viitor conditional)))
|
||||
(compound-past "perfectul compus = a-avea-clitic + INVARIABLE participle")
|
||||
(perfect-aux "a avea (am/ai/a/am/ați/au) — ONE auxiliary for ALL verbs")
|
||||
(perfect-aux-split no) ; << SIMPLER than Italian: no essere/avere selection
|
||||
(participle-agreement none) ; invariable in the perfect compus (agrees only as
|
||||
; an adjective / in the passive)
|
||||
(future "voi/vei/va/vom/veți/vor + infinitive (viitor literar)")
|
||||
(conditional "aș/ai/ar/am/ați/ar + infinitive")
|
||||
(subjunctive "conjunctiv: particle 'să' + subjunctive present")
|
||||
(modal-complement "modal + să + subjunctive (vreau să merg, poți să ajuți)")
|
||||
(copula "a fi")
|
||||
(passive "a fi + participle (participle AGREES like an adjective)")
|
||||
(comparative "mai / mai puțin ADJ decât")
|
||||
|
||||
;; ── clitic system (partial — see honest gaps) ───────────────────────────
|
||||
(clitics yes)
|
||||
(clitic-set ((acc mă te îl o ne vă îi le) (dat îmi îți îi ne vă le)
|
||||
(refl mă te se ne vă se)))
|
||||
(clitic-placement ((finite proclitic) ; îmi place, o văd
|
||||
(perfect-compus elision) ; << m-am, l-am, i-am (PARTIAL)
|
||||
(imperative-affirmative enclitic))) ; dă-mi (PARTIAL)
|
||||
|
||||
;; ── SACRED safety bar (shared with es/pt/it/en) ─────────────────────────
|
||||
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
|
||||
(negation "nu (single preverbal marker) + concord")
|
||||
(negative-concord yes) ; nu … nimic / nimeni / niciodată / niciun
|
||||
(negative-imperative "nu + INFINITIVE : nu pleca! (KNOWN GAP: uses imperative stem)"))
|
||||
@@ -250,6 +250,7 @@ fn en_irregular_verb(base: String) -> [String] {
|
||||
if str_eq(base, "cut") { let r: [String] = ["cut", "cuts", "cut", "cut", "cutting"]; return r }
|
||||
if str_eq(base, "set") { let r: [String] = ["set", "sets", "set", "set", "setting"]; return r }
|
||||
if str_eq(base, "hit") { let r: [String] = ["hit", "hits", "hit", "hit", "hitting"]; return r }
|
||||
if str_eq(base, "fight") { let r: [String] = ["fight", "fights","fought", "fought", "fighting"]; return r }
|
||||
return empty
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,280 @@
|
||||
// multilingual.el - the language layer for the native-el interlocutor.
|
||||
//
|
||||
// Deterministic, NO generative model (ports multilingual.py):
|
||||
// 1. ml_detect(text) -> ISO code (en/es/pt/it) via stopword + diacritic score
|
||||
// 2. ml_tr(key, lang) -> localized fixed phrase (SACRED per-language yes/no/decline)
|
||||
// 3. ml_term(w, lang) -> PT/ES content term -> EN engram equivalent
|
||||
// 4. ml_translate_pred(lemma, lang) -> EN predicate lemma -> target infinitive
|
||||
//
|
||||
// The Python detector count-weights stopwords and diacritics; here diacritics are
|
||||
// scored by PRESENCE (str_contains) rather than codepoint counting, to stay clear
|
||||
// of UTF-8 index hazards in the runtime. Faithful enough to classify typical
|
||||
// queries; documented simplification. Depends on: comprehend (cp_tokenize).
|
||||
|
||||
// ── 1. language detection ─────────────────────────────────────────────────────
|
||||
|
||||
fn ml_stop_en(w: String) -> Bool {
|
||||
if str_eq(w, "the") { return true }
|
||||
if str_eq(w, "does") { return true }
|
||||
if str_eq(w, "do") { return true }
|
||||
if str_eq(w, "did") { return true }
|
||||
if str_eq(w, "what") { return true }
|
||||
if str_eq(w, "who") { return true }
|
||||
if str_eq(w, "is") { return true }
|
||||
if str_eq(w, "are") { return true }
|
||||
if str_eq(w, "how") { return true }
|
||||
if str_eq(w, "you") { return true }
|
||||
if str_eq(w, "your") { return true }
|
||||
if str_eq(w, "of") { return true }
|
||||
if str_eq(w, "to") { return true }
|
||||
if str_eq(w, "and") { return true }
|
||||
if str_eq(w, "for") { return true }
|
||||
if str_eq(w, "explain") { return true }
|
||||
if str_eq(w, "answer") { return true }
|
||||
if str_eq(w, "memory") { return true }
|
||||
if str_eq(w, "with") { return true }
|
||||
if str_eq(w, "not") { return true }
|
||||
if str_eq(w, "store") { return true }
|
||||
return false
|
||||
}
|
||||
|
||||
fn ml_stop_es(w: String) -> Bool {
|
||||
if str_eq(w, "que") { return true }
|
||||
if str_eq(w, "qué") { return true }
|
||||
if str_eq(w, "una") { return true }
|
||||
if str_eq(w, "usted") { return true }
|
||||
if str_eq(w, "su") { return true }
|
||||
if str_eq(w, "cómo") { return true }
|
||||
if str_eq(w, "como") { return true }
|
||||
if str_eq(w, "cuál") { return true }
|
||||
if str_eq(w, "quién") { return true }
|
||||
if str_eq(w, "está") { return true }
|
||||
if str_eq(w, "es") { return true }
|
||||
if str_eq(w, "los") { return true }
|
||||
if str_eq(w, "las") { return true }
|
||||
if str_eq(w, "del") { return true }
|
||||
if str_eq(w, "al") { return true }
|
||||
if str_eq(w, "explica") { return true }
|
||||
if str_eq(w, "explique") { return true }
|
||||
if str_eq(w, "forma") { return true }
|
||||
if str_eq(w, "con") { return true }
|
||||
if str_eq(w, "memoria") { return true }
|
||||
if str_eq(w, "responde") { return true }
|
||||
return false
|
||||
}
|
||||
|
||||
fn ml_stop_pt(w: String) -> Bool {
|
||||
if str_eq(w, "que") { return true }
|
||||
if str_eq(w, "uma") { return true }
|
||||
if str_eq(w, "você") { return true }
|
||||
if str_eq(w, "sua") { return true }
|
||||
if str_eq(w, "seu") { return true }
|
||||
if str_eq(w, "como") { return true }
|
||||
if str_eq(w, "memória") { return true }
|
||||
if str_eq(w, "isso") { return true }
|
||||
if str_eq(w, "os") { return true }
|
||||
if str_eq(w, "as") { return true }
|
||||
if str_eq(w, "da") { return true }
|
||||
if str_eq(w, "do") { return true }
|
||||
if str_eq(w, "na") { return true }
|
||||
if str_eq(w, "no") { return true }
|
||||
if str_eq(w, "explica") { return true }
|
||||
if str_eq(w, "forma") { return true }
|
||||
if str_eq(w, "é") { return true }
|
||||
if str_eq(w, "está") { return true }
|
||||
if str_eq(w, "com") { return true }
|
||||
if str_eq(w, "responda") { return true }
|
||||
return false
|
||||
}
|
||||
|
||||
fn ml_stop_it(w: String) -> Bool {
|
||||
if str_eq(w, "che") { return true }
|
||||
if str_eq(w, "una") { return true }
|
||||
if str_eq(w, "come") { return true }
|
||||
if str_eq(w, "della") { return true }
|
||||
if str_eq(w, "gli") { return true }
|
||||
if str_eq(w, "è") { return true }
|
||||
if str_eq(w, "sono") { return true }
|
||||
if str_eq(w, "questo") { return true }
|
||||
if str_eq(w, "nel") { return true }
|
||||
if str_eq(w, "di") { return true }
|
||||
if str_eq(w, "il") { return true }
|
||||
if str_eq(w, "cosa") { return true }
|
||||
if str_eq(w, "per") { return true }
|
||||
if str_eq(w, "memoria") { return true }
|
||||
if str_eq(w, "spiega") { return true }
|
||||
if str_eq(w, "rispondi") { return true }
|
||||
return false
|
||||
}
|
||||
|
||||
// diacritic PRESENCE score (weight 3 each; hard overrides weight 8).
|
||||
fn ml_dia_score(low: String, lang: String) -> Int {
|
||||
let s: Int = 0
|
||||
if str_eq(lang, "pt") {
|
||||
if str_contains(low, "ã") { let s = s + 3 }
|
||||
if str_contains(low, "õ") { let s = s + 3 }
|
||||
if str_contains(low, "ç") { let s = s + 3 }
|
||||
if str_contains(low, "ê") { let s = s + 3 }
|
||||
if str_contains(low, "á") { let s = s + 3 }
|
||||
// hard PT markers (ã/õ almost never appear outside PT)
|
||||
if str_contains(low, "ã") { let s = s + 8 }
|
||||
if str_contains(low, "õ") { let s = s + 8 }
|
||||
}
|
||||
if str_eq(lang, "es") {
|
||||
if str_contains(low, "ñ") { let s = s + 3 }
|
||||
if str_contains(low, "¿") { let s = s + 3 }
|
||||
if str_contains(low, "¡") { let s = s + 3 }
|
||||
if str_contains(low, "á") { let s = s + 3 }
|
||||
if str_contains(low, "é") { let s = s + 3 }
|
||||
// hard ES markers
|
||||
if str_contains(low, "ñ") { let s = s + 8 }
|
||||
if str_contains(low, "¿") { let s = s + 8 }
|
||||
if str_contains(low, "¡") { let s = s + 8 }
|
||||
}
|
||||
if str_eq(lang, "it") {
|
||||
if str_contains(low, "è") { let s = s + 3 }
|
||||
if str_contains(low, "ì") { let s = s + 3 }
|
||||
if str_contains(low, "ò") { let s = s + 3 }
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
fn ml_stop_score(toks: [String], lang: String) -> Int {
|
||||
let n: Int = native_list_len(toks)
|
||||
let s: Int = 0
|
||||
let i: Int = 0
|
||||
while i < n {
|
||||
let w: String = native_list_get(toks, i)
|
||||
if str_eq(lang, "en") { if ml_stop_en(w) { let s = s + 2 } }
|
||||
if str_eq(lang, "es") { if ml_stop_es(w) { let s = s + 2 } }
|
||||
if str_eq(lang, "pt") { if ml_stop_pt(w) { let s = s + 2 } }
|
||||
if str_eq(lang, "it") { if ml_stop_it(w) { let s = s + 2 } }
|
||||
let i = i + 1
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
fn ml_detect(text: String) -> String {
|
||||
if str_eq(text, "") { return "en" }
|
||||
let low: String = str_to_lower(text)
|
||||
let toks: [String] = cp_tokenize(text)
|
||||
// NOTE: el's overloaded `+` mis-compiles two chained function-call Int operands
|
||||
// as string concat (documented in comprehend_gate.el). Bind each call to an Int
|
||||
// var and add vars one at a time so the addition stays integer.
|
||||
let en: Int = ml_stop_score(toks, "en")
|
||||
let es_s: Int = ml_stop_score(toks, "es")
|
||||
let es_d: Int = ml_dia_score(low, "es")
|
||||
let es: Int = es_s + es_d
|
||||
let pt_s: Int = ml_stop_score(toks, "pt")
|
||||
let pt_d: Int = ml_dia_score(low, "pt")
|
||||
let pt: Int = pt_s + pt_d
|
||||
let it_s: Int = ml_stop_score(toks, "it")
|
||||
let it_d: Int = ml_dia_score(low, "it")
|
||||
let it: Int = it_s + it_d
|
||||
|
||||
let best: String = "en"
|
||||
let bs: Int = en
|
||||
if es > bs { let best = "es"; let bs = es }
|
||||
if pt > bs { let best = "pt"; let bs = pt }
|
||||
if it > bs { let best = "it"; let bs = it }
|
||||
// weak signal -> honest fallback to English
|
||||
if bs < 3 { return "en" }
|
||||
return best
|
||||
}
|
||||
|
||||
// ── 2. localized fixed phrases (SACRED per-language decline/yes/no) ────────────
|
||||
|
||||
fn ml_tr(key: String, lang: String) -> String {
|
||||
if str_eq(key, "no_memory") {
|
||||
if str_eq(lang, "pt") { return "Não tenho isso na minha memória." }
|
||||
if str_eq(lang, "es") { return "No tengo eso en mi memoria." }
|
||||
if str_eq(lang, "it") { return "Non ho quello nella mia memoria." }
|
||||
return "I don't have that in my memory."
|
||||
}
|
||||
if str_eq(key, "parse_fail") {
|
||||
if str_eq(lang, "pt") { return "Não consegui interpretar isso." }
|
||||
if str_eq(lang, "es") { return "No pude interpretar eso." }
|
||||
if str_eq(lang, "it") { return "Non sono riuscito a interpretarlo." }
|
||||
return "I didn't parse that."
|
||||
}
|
||||
if str_eq(key, "yes") {
|
||||
if str_eq(lang, "pt") { return "Sim" }
|
||||
if str_eq(lang, "es") { return "Sí" }
|
||||
if str_eq(lang, "it") { return "Sì" }
|
||||
return "Yes"
|
||||
}
|
||||
if str_eq(key, "no") {
|
||||
if str_eq(lang, "pt") { return "Não" }
|
||||
if str_eq(lang, "es") { return "No" }
|
||||
if str_eq(lang, "it") { return "No" }
|
||||
return "No"
|
||||
}
|
||||
if str_eq(key, "identity") {
|
||||
if str_eq(lang, "pt") { return "Sou o Neuron, o engrama com quem você está falando." }
|
||||
if str_eq(lang, "es") { return "Soy Neuron, el engrama con el que estás hablando." }
|
||||
if str_eq(lang, "it") { return "Sono Neuron, l'engramma con cui stai parlando." }
|
||||
return "I'm Neuron, the engram you're speaking with."
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// ── 3. retrieval term lexicon (PT/ES content term -> EN engram equivalent) ─────
|
||||
|
||||
fn ml_term(w: String, lang: String) -> String {
|
||||
if str_eq(lang, "en") { return w }
|
||||
if str_eq(w, "saliência") { return "salience" }
|
||||
if str_eq(w, "saliencia") { return "salience" }
|
||||
if str_eq(w, "memória") { return "memory" }
|
||||
if str_eq(w, "memoria") { return "memory" }
|
||||
if str_eq(w, "geometria") { return "geometry" }
|
||||
if str_eq(w, "geometrias") { return "geometry" }
|
||||
if str_eq(w, "geometrías") { return "geometry" }
|
||||
if str_eq(w, "forma") { return "form" }
|
||||
if str_eq(w, "consolidação") { return "consolidation" }
|
||||
if str_eq(w, "consolidación") { return "consolidation" }
|
||||
if str_eq(w, "aprendizagem") { return "learning" }
|
||||
if str_eq(w, "aprendizaje") { return "learning" }
|
||||
if str_eq(w, "nó") { return "node" }
|
||||
if str_eq(w, "nodo") { return "node" }
|
||||
if str_eq(w, "armazenamento") { return "storage" }
|
||||
if str_eq(w, "almacenamiento") { return "storage" }
|
||||
if str_eq(w, "estrutura") { return "structure" }
|
||||
if str_eq(w, "estructura") { return "structure" }
|
||||
return w
|
||||
}
|
||||
|
||||
// ── 4. predicate translation (EN lemma -> target infinitive; pass-through) ─────
|
||||
|
||||
fn ml_translate_pred(lemma: String, lang: String) -> String {
|
||||
if str_eq(lang, "en") { return lemma }
|
||||
if str_eq(lang, "es") {
|
||||
if str_eq(lemma, "store") { return "almacenar" }
|
||||
if str_eq(lemma, "use") { return "usar" }
|
||||
if str_eq(lemma, "have") { return "tener" }
|
||||
if str_eq(lemma, "be") { return "ser" }
|
||||
if str_eq(lemma, "give") { return "dar" }
|
||||
if str_eq(lemma, "make") { return "hacer" }
|
||||
if str_eq(lemma, "learn") { return "aprender" }
|
||||
if str_eq(lemma, "form") { return "formar" }
|
||||
return lemma
|
||||
}
|
||||
if str_eq(lang, "pt") {
|
||||
if str_eq(lemma, "store") { return "armazenar" }
|
||||
if str_eq(lemma, "use") { return "usar" }
|
||||
if str_eq(lemma, "have") { return "ter" }
|
||||
if str_eq(lemma, "be") { return "ser" }
|
||||
if str_eq(lemma, "give") { return "dar" }
|
||||
if str_eq(lemma, "make") { return "fazer" }
|
||||
if str_eq(lemma, "learn") { return "aprender" }
|
||||
if str_eq(lemma, "form") { return "formar" }
|
||||
return lemma
|
||||
}
|
||||
if str_eq(lang, "it") {
|
||||
if str_eq(lemma, "store") { return "memorizzare" }
|
||||
if str_eq(lemma, "use") { return "usare" }
|
||||
if str_eq(lemma, "have") { return "avere" }
|
||||
if str_eq(lemma, "be") { return "essere" }
|
||||
return lemma
|
||||
}
|
||||
return lemma
|
||||
}
|
||||
@@ -0,0 +1,140 @@
|
||||
// propositions.el - the READ primitive over the engram's OWN memories, native el.
|
||||
//
|
||||
// Free memory text -> structured PROPOSITIONS (triples):
|
||||
// (subject, predicate, object, modifiers, polarity, tense, source, confidence)
|
||||
//
|
||||
// This is comprehension turned inward: the Python reference (propositions.py) ran
|
||||
// spaCy's dependency parser over each memory sentence and walked the arcs. Here
|
||||
// the spaCy role is filled by the el-native parser (comprehend.el / parse_spec):
|
||||
// each sentence is parsed to a meaning-spec, and the spec's roles ARE the triple.
|
||||
// Nothing generates text. NEGATION IS SACRED: polarity flows straight from the
|
||||
// spec's polarity field and is never dropped or inverted.
|
||||
//
|
||||
// Depends on: comprehend (parse_spec / parse_spec_lang), grammar (slots_get).
|
||||
|
||||
// ── sentence segmentation ─────────────────────────────────────────────────────
|
||||
// Split on sentence-final punctuation (. ! ?) and hard newlines. Markdown/long
|
||||
// memories are handled shallowly (the reference caps + ranks by query overlap;
|
||||
// that ranking belongs to the dialogue layer, not here).
|
||||
|
||||
fn prop_is_boundary(c: String) -> Bool {
|
||||
if str_eq(c, ".") { return true }
|
||||
if str_eq(c, "!") { return true }
|
||||
if str_eq(c, "?") { return true }
|
||||
if str_eq(c, "\n") { return true }
|
||||
return false
|
||||
}
|
||||
|
||||
fn prop_split_sentences(text: String) -> [String] {
|
||||
let out: [String] = native_list_empty()
|
||||
let n: Int = str_len(text)
|
||||
let start: Int = 0
|
||||
let i: Int = 0
|
||||
while i < n {
|
||||
let c: String = str_slice(text, i, i + 1)
|
||||
if prop_is_boundary(c) {
|
||||
let seg: String = str_slice(text, start, i + 1)
|
||||
let trimmed: String = cp_trim_punct(seg)
|
||||
if !str_eq(trimmed, "") {
|
||||
let out = native_list_append(out, seg)
|
||||
}
|
||||
let start = i + 1
|
||||
}
|
||||
let i = i + 1
|
||||
}
|
||||
if start < n {
|
||||
let seg: String = str_slice(text, start, n)
|
||||
let trimmed: String = cp_trim_punct(seg)
|
||||
if !str_eq(trimmed, "") {
|
||||
let out = native_list_append(out, seg)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// ── spec -> proposition record ────────────────────────────────────────────────
|
||||
// A proposition is a slot map (same [String] shape as the spec) with the READ
|
||||
// contract keys. Modifiers fold the spec's location + iobj adjuncts.
|
||||
|
||||
fn prop_confidence(subject: String, predicate: String, object: String) -> String {
|
||||
if str_eq(predicate, "") { return "0.0" }
|
||||
if str_eq(subject, "") { return "0.4" }
|
||||
if str_eq(object, "") { return "0.7" }
|
||||
return "1.0"
|
||||
}
|
||||
|
||||
fn prop_modifiers(spec: [String]) -> String {
|
||||
let loc: String = slots_get(spec, "location")
|
||||
let iobj: String = slots_get(spec, "iobj")
|
||||
let parts: [String] = native_list_empty()
|
||||
if !str_eq(loc, "") { let parts = native_list_append(parts, loc) }
|
||||
if !str_eq(iobj, "") { let parts = native_list_append(parts, "to " + iobj) }
|
||||
return str_join(parts, "; ")
|
||||
}
|
||||
|
||||
fn prop_from_spec(spec: [String], source_id: String) -> [String] {
|
||||
let subject: String = slots_get(spec, "agent")
|
||||
let predicate: String = slots_get(spec, "predicate")
|
||||
let object: String = slots_get(spec, "patient")
|
||||
let polarity: String = slots_get(spec, "polarity")
|
||||
let tense: String = slots_get(spec, "tense")
|
||||
let mods: String = prop_modifiers(spec)
|
||||
let conf: String = prop_confidence(subject, predicate, object)
|
||||
|
||||
let p: [String] = native_list_empty()
|
||||
let p = native_list_append(p, "subject"); let p = native_list_append(p, subject)
|
||||
let p = native_list_append(p, "predicate"); let p = native_list_append(p, predicate)
|
||||
let p = native_list_append(p, "object"); let p = native_list_append(p, object)
|
||||
let p = native_list_append(p, "modifiers"); let p = native_list_append(p, mods)
|
||||
let p = native_list_append(p, "polarity"); let p = native_list_append(p, polarity)
|
||||
let p = native_list_append(p, "tense"); let p = native_list_append(p, tense)
|
||||
let p = native_list_append(p, "source"); let p = native_list_append(p, source_id)
|
||||
let p = native_list_append(p, "confidence"); let p = native_list_append(p, conf)
|
||||
return p
|
||||
}
|
||||
|
||||
// Extract one proposition from a single sentence (given language).
|
||||
fn prop_extract_one_lang(sentence: String, lang: String, source_id: String) -> [String] {
|
||||
let spec: [String] = parse_spec_lang(sentence, lang)
|
||||
return prop_from_spec(spec, source_id)
|
||||
}
|
||||
|
||||
fn prop_extract_one(sentence: String, source_id: String) -> [String] {
|
||||
return prop_extract_one_lang(sentence, "en", source_id)
|
||||
}
|
||||
|
||||
// Render a proposition as a compact trace line (repr parity with propositions.py).
|
||||
fn prop_repr(p: [String]) -> String {
|
||||
let neg: String = ""
|
||||
if str_eq(slots_get(p, "polarity"), "neg") { let neg = "NOT " }
|
||||
let mods: String = slots_get(p, "modifiers")
|
||||
let modstr: String = ""
|
||||
if !str_eq(mods, "") { let modstr = " [" + mods + "]" }
|
||||
let s: String = "(" + slots_get(p, "subject") + " -" + neg + slots_get(p, "predicate")
|
||||
let s = s + "-> " + slots_get(p, "object") + modstr
|
||||
let s = s + " conf=" + slots_get(p, "confidence") + ")"
|
||||
return s
|
||||
}
|
||||
|
||||
// Extract all propositions from a memory's text (one per sentence). Returns a
|
||||
// flat [String] whose entries are the prop_repr trace lines, in reading order.
|
||||
fn prop_extract_lang(text: String, lang: String, source_id: String) -> [String] {
|
||||
let sents: [String] = prop_split_sentences(text)
|
||||
let m: Int = native_list_len(sents)
|
||||
let out: [String] = native_list_empty()
|
||||
let i: Int = 0
|
||||
while i < m {
|
||||
let sent: String = native_list_get(sents, i)
|
||||
let p: [String] = prop_extract_one_lang(sent, lang, source_id)
|
||||
// drop empty parses (no predicate recovered): honest partial, not noise.
|
||||
if !str_eq(slots_get(p, "predicate"), "") {
|
||||
let out = native_list_append(out, prop_repr(p))
|
||||
}
|
||||
let i = i + 1
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
fn prop_extract(text: String, source_id: String) -> [String] {
|
||||
return prop_extract_lang(text, "en", source_id)
|
||||
}
|
||||
@@ -248,6 +248,56 @@ fn add_punct(s: String, intent: String) -> String {
|
||||
return s + "."
|
||||
}
|
||||
|
||||
// ── Polarity-aware negation (SACRED field honored on the generation side) ─────
|
||||
//
|
||||
// Negation must never be dropped between comprehension and realization. The
|
||||
// meaning-spec carries an explicit "polarity" field ("aff"|"neg") and optional
|
||||
// "neg_word" (standalone negative adverb, e.g. "never"). English uses
|
||||
// do-support ("did not see") or preverbal adverb ("never fought"); copular "be"
|
||||
// takes post-verbal "not"; other languages get a preverbal negator particle.
|
||||
|
||||
fn realize_negator(code: String) -> String {
|
||||
if str_eq(code, "es") { return "no" }
|
||||
if str_eq(code, "pt") { return "não" }
|
||||
if str_eq(code, "ca") { return "no" }
|
||||
if str_eq(code, "it") { return "non" }
|
||||
if str_eq(code, "fr") { return "ne" }
|
||||
if str_eq(code, "de") { return "nicht" }
|
||||
if str_eq(code, "ro") { return "nu" }
|
||||
return "not"
|
||||
}
|
||||
|
||||
fn realize_assert_neg_en(predicate: String, tense: String, person: String, number: String, agent: String, patient: String, iobj: String, location: String, neg_word: String, profile: [String]) -> String {
|
||||
let parts: [String] = native_list_empty()
|
||||
let parts = native_list_append(parts, agent)
|
||||
if !str_eq(neg_word, "") {
|
||||
// adverbial negation: "I never fought the ocean."
|
||||
let verb_surf: String = morph_conjugate(predicate, tense, person, number, profile)
|
||||
let parts = native_list_append(parts, neg_word)
|
||||
let parts = native_list_append(parts, verb_surf)
|
||||
} else {
|
||||
if str_eq(predicate, "be") {
|
||||
// copular: "she was not a monster"
|
||||
let be_form: String = morph_conjugate("be", tense, person, number, profile)
|
||||
let parts = native_list_append(parts, be_form)
|
||||
let parts = native_list_append(parts, "not")
|
||||
} else {
|
||||
// do-support: "she did not see the man"
|
||||
let do_form: String = morph_conjugate("do", tense, person, number, profile)
|
||||
let parts = native_list_append(parts, do_form)
|
||||
let parts = native_list_append(parts, "not")
|
||||
let parts = native_list_append(parts, predicate)
|
||||
}
|
||||
}
|
||||
if !str_eq(patient, "") { let parts = native_list_append(parts, patient) }
|
||||
if !str_eq(iobj, "") {
|
||||
let parts = native_list_append(parts, "to")
|
||||
let parts = native_list_append(parts, iobj)
|
||||
}
|
||||
if !str_eq(location, "") { let parts = native_list_append(parts, location) }
|
||||
return str_join(parts, " ")
|
||||
}
|
||||
|
||||
// ── Main realization entry point ──────────────────────────────────────────────
|
||||
|
||||
fn realize_lang(form: [String], profile: [String]) -> String {
|
||||
@@ -284,6 +334,50 @@ fn realize_lang(form: [String], profile: [String]) -> String {
|
||||
}
|
||||
|
||||
// ── Assertion (declarative) ───────────────────────────────────────────────
|
||||
let polarity: String = slots_get(form, "polarity")
|
||||
let neg_word: String = slots_get(form, "neg_word")
|
||||
let iobj: String = slots_get(form, "iobj")
|
||||
let code: String = lang_get(profile, "code")
|
||||
|
||||
// Subordinate clause tail (SACRED completeness — the clause is carried, never
|
||||
// dropped): "<conj> <subordinate surface>", e.g. "because he was a monster".
|
||||
let subord_conj: String = slots_get(form, "subord_conj")
|
||||
let subord_text: String = slots_get(form, "subord_text")
|
||||
let subord_tail: String = ""
|
||||
if !str_eq(subord_conj, "") {
|
||||
if !str_eq(subord_text, "") {
|
||||
let subord_tail = subord_conj + " " + subord_text
|
||||
} else {
|
||||
let subord_tail = subord_conj
|
||||
}
|
||||
}
|
||||
|
||||
// Negative polarity: SACRED — never dropped.
|
||||
if str_eq(polarity, "neg") {
|
||||
if str_eq(code, "en") {
|
||||
let sentence: String = realize_assert_neg_en(predicate, tense, person, number, agent, patient, iobj, location, neg_word, profile)
|
||||
return add_punct(capitalize_first(sentence), "assert")
|
||||
}
|
||||
// Generic non-English: affirmative core with a preverbal negator particle.
|
||||
let neg_particle: String = realize_negator(code)
|
||||
let vp_pair: [String] = realize_vp_lang(predicate, tense, aspect, person, number, profile)
|
||||
let verb_surf: String = native_list_get(vp_pair, 0)
|
||||
let aux_surf: String = native_list_get(vp_pair, 1)
|
||||
let vp_str: String = neg_particle + " " + gram_build_vp(verb_surf, aux_surf, profile)
|
||||
let core: String = gram_order_constituents(agent, vp_str, patient, profile)
|
||||
let parts: [String] = native_list_empty()
|
||||
let parts = native_list_append(parts, core)
|
||||
if !str_eq(iobj, "") {
|
||||
let parts = native_list_append(parts, "to")
|
||||
let parts = native_list_append(parts, iobj)
|
||||
}
|
||||
if !str_eq(location, "") { let parts = native_list_append(parts, location) }
|
||||
if !str_eq(subord_tail, "") { let parts = native_list_append(parts, subord_tail) }
|
||||
let sentence: String = str_join(parts, " ")
|
||||
return add_punct(capitalize_first(sentence), "assert")
|
||||
}
|
||||
|
||||
// Affirmative.
|
||||
let vp_pair: [String] = realize_vp_lang(predicate, tense, aspect, person, number, profile)
|
||||
let verb_surf: String = native_list_get(vp_pair, 0)
|
||||
let aux_surf: String = native_list_get(vp_pair, 1)
|
||||
@@ -293,9 +387,16 @@ fn realize_lang(form: [String], profile: [String]) -> String {
|
||||
|
||||
let parts: [String] = native_list_empty()
|
||||
let parts = native_list_append(parts, core)
|
||||
if !str_eq(iobj, "") {
|
||||
let parts = native_list_append(parts, "to")
|
||||
let parts = native_list_append(parts, iobj)
|
||||
}
|
||||
if !str_eq(location, "") {
|
||||
let parts = native_list_append(parts, location)
|
||||
}
|
||||
if !str_eq(subord_tail, "") {
|
||||
let parts = native_list_append(parts, subord_tail)
|
||||
}
|
||||
let sentence: String = str_join(parts, " ")
|
||||
return add_punct(capitalize_first(sentence), "assert")
|
||||
}
|
||||
|
||||
@@ -0,0 +1,180 @@
|
||||
// self_region.el — the engram's REAL self/identity region, pulled at query time
|
||||
// (native el). This replaces the hardcoded identity anchors and the canned
|
||||
// "I'm Neuron, the engram you're speaking with." template: the identity LANDING
|
||||
// signal and the identity READOUT both come from the engram's own Self/identity
|
||||
// nodes, read through the in-process engram el API.
|
||||
//
|
||||
// Port of self_region.py. The Python module precomputed MiniLM landing vectors;
|
||||
// here the engram's own store IS the geometry — we pull the self nodes by
|
||||
// single-term lexical search (the engram search is a single-term matcher, so we
|
||||
// pool several probes) and rank them by self-signal. No text is generated; the
|
||||
// readout is the self nodes' OWN prose, verbatim (SACRED negation survives by
|
||||
// construction — we never paraphrase, so a negated self-statement stays negated).
|
||||
//
|
||||
// ENGRAM el API NOTE: engram_search_json / engram_get_node_json / engram_node_full
|
||||
// / engram_connect are C runtime builtins. Their argument order is the C order
|
||||
// (engram_connect(from, to, weight, relation)), NOT the runtime/engram.el wrapper
|
||||
// order — we call the builtins directly and never concatenate that wrapper.
|
||||
//
|
||||
// Depends on: comprehend (str helpers via runtime), propositions (prop_split_sentences),
|
||||
// multilingual (ml_tr), the engram builtins, the json builtins.
|
||||
|
||||
// ── single-term self probes (pooled, because engram search is single-term) ────
|
||||
fn sr_terms() -> [String] {
|
||||
let t: [String] = native_list_empty()
|
||||
let t = native_list_append(t, "self")
|
||||
let t = native_list_append(t, "identity")
|
||||
let t = native_list_append(t, "Neuron")
|
||||
let t = native_list_append(t, "consciousness")
|
||||
let t = native_list_append(t, "values")
|
||||
let t = native_list_append(t, "continuous")
|
||||
return t
|
||||
}
|
||||
|
||||
// The canonical self-root: content begins "# self" or label is "# self"/"self".
|
||||
fn sr_is_root(content: String, label: String) -> Bool {
|
||||
let lc: String = str_to_lower(content)
|
||||
let ll: String = str_to_lower(str_trim(label))
|
||||
if str_starts_with(lc, "# self") { return true }
|
||||
if str_eq(ll, "# self") { return true }
|
||||
if str_eq(ll, "self") { return true }
|
||||
return false
|
||||
}
|
||||
|
||||
// How strongly a node belongs to the self/identity region (integer points, to
|
||||
// avoid el's float-in-`+` pitfalls). Mirrors _self_score in self_region.py.
|
||||
fn sr_score(node_json: String) -> Int {
|
||||
let content: String = json_get_string(node_json, "content")
|
||||
let label: String = json_get_string(node_json, "label")
|
||||
let tags: String = str_to_lower(json_get_string(node_json, "tags"))
|
||||
let low: String = str_to_lower(content)
|
||||
let s: Int = 0
|
||||
// identity tags
|
||||
if str_contains(tags, "self") { let s = s + 2 }
|
||||
if str_contains(tags, "identity") { let s = s + 2 }
|
||||
if str_contains(tags, "self-model") { let s = s + 2 }
|
||||
if str_contains(tags, "consciousness") { let s = s + 2 }
|
||||
if str_contains(tags, "memory-philosophy") { let s = s + 2 }
|
||||
// the named self-traversal root
|
||||
if sr_is_root(content, label) { let s = s + 12 }
|
||||
if str_contains(low, "who i am") { let s = s + 3 }
|
||||
if str_contains(low, "i am neuron") { let s = s + 3 }
|
||||
// softer identity keywords
|
||||
if str_contains(low, "my values") { let s = s + 1 }
|
||||
if str_contains(low, "my purpose") { let s = s + 1 }
|
||||
if str_contains(low, "identity") { let s = s + 1 }
|
||||
return s
|
||||
}
|
||||
|
||||
// list-contains helper (dedup self-node ids across the pooled probes).
|
||||
fn sr_ids_has(ids: [String], id: String) -> Bool {
|
||||
let n: Int = native_list_len(ids)
|
||||
let i: Int = 0
|
||||
while i < n {
|
||||
if str_eq(native_list_get(ids, i), id) { return true }
|
||||
let i = i + 1
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// Pull the self nodes: pool every probe's hits, dedupe by id, keep only nodes
|
||||
// with genuine self-signal (score >= 1). Returns the node-json strings.
|
||||
fn sr_pull() -> [String] {
|
||||
let terms: [String] = sr_terms()
|
||||
let nt: Int = native_list_len(terms)
|
||||
let seen: [String] = native_list_empty()
|
||||
let out: [String] = native_list_empty()
|
||||
let ti: Int = 0
|
||||
while ti < nt {
|
||||
let term: String = native_list_get(terms, ti)
|
||||
let hits: String = engram_search_json(term, 30)
|
||||
let hn: Int = json_array_len(hits)
|
||||
let hi: Int = 0
|
||||
while hi < hn {
|
||||
let node: String = json_array_get(hits, hi)
|
||||
let id: String = json_get_string(node, "id")
|
||||
if !str_eq(id, "") {
|
||||
if !sr_ids_has(seen, id) {
|
||||
let seen = native_list_append(seen, id)
|
||||
if sr_score(node) >= 1 {
|
||||
let out = native_list_append(out, node)
|
||||
}
|
||||
}
|
||||
}
|
||||
let hi = hi + 1
|
||||
}
|
||||
let ti = ti + 1
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// Return the single highest-signal self node (the readout seed), or "" if the
|
||||
// self region is thin/empty. We keep it O(n) — pick the max-score node, with the
|
||||
// canonical root strongly favored by sr_score's +12.
|
||||
fn sr_best_node() -> String {
|
||||
let nodes: [String] = sr_pull()
|
||||
let n: Int = native_list_len(nodes)
|
||||
let best: String = ""
|
||||
let best_s: Int = 0
|
||||
let i: Int = 0
|
||||
while i < n {
|
||||
let node: String = native_list_get(nodes, i)
|
||||
let s: Int = sr_score(node)
|
||||
if s > best_s {
|
||||
let best_s = s
|
||||
let best = node
|
||||
}
|
||||
let i = i + 1
|
||||
}
|
||||
return best
|
||||
}
|
||||
|
||||
fn sr_available() -> Bool {
|
||||
if str_eq(sr_best_node(), "") { return false }
|
||||
return true
|
||||
}
|
||||
|
||||
// Read out the identity from the REAL self node: lead with the first first-person
|
||||
// self-statement ("I am Neuron …"), then one more grounded self line if present.
|
||||
// Verbatim from the node's own prose — no template, negation SACRED. Falls back
|
||||
// to the localized identity phrase ONLY if the live pull is empty (logged shape).
|
||||
fn sr_readout(lang: String) -> String {
|
||||
let node: String = sr_best_node()
|
||||
if str_eq(node, "") {
|
||||
// honest fallback — the self region is unreachable/thin.
|
||||
return ml_tr("identity", lang)
|
||||
}
|
||||
let content: String = json_get_string(node, "content")
|
||||
let sents: [String] = prop_split_sentences(content)
|
||||
let ns: Int = native_list_len(sents)
|
||||
let lead: String = ""
|
||||
let second: String = ""
|
||||
let i: Int = 0
|
||||
while i < ns {
|
||||
let raw: String = str_trim(native_list_get(sents, i))
|
||||
// strip a leading markdown heading marker
|
||||
let s: String = raw
|
||||
if str_starts_with(s, "# ") { let s = str_trim(str_slice(s, 2, str_len(s))) }
|
||||
let low: String = str_to_lower(s)
|
||||
let is_fp: Bool = false
|
||||
if str_starts_with(s, "I ") { let is_fp = true }
|
||||
if str_starts_with(s, "I'm") { let is_fp = true }
|
||||
if str_contains(low, "i am neuron") { let is_fp = true }
|
||||
if is_fp {
|
||||
if str_eq(lead, "") {
|
||||
let lead = s
|
||||
} else {
|
||||
if str_eq(second, "") { let second = s }
|
||||
}
|
||||
}
|
||||
let i = i + 1
|
||||
}
|
||||
if str_eq(lead, "") {
|
||||
// no first-person line — read out the first non-empty sentence verbatim.
|
||||
if ns > 0 { let lead = str_trim(native_list_get(sents, 0)) }
|
||||
}
|
||||
if str_eq(lead, "") { return ml_tr("identity", lang) }
|
||||
let out: String = lead
|
||||
if !str_eq(second, "") { let out = out + " " + second }
|
||||
return out
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
+144861
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
+130676
File diff suppressed because it is too large
Load Diff
+193894
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
+115916
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,93 @@
|
||||
// comprehend_gate.el - the TELEPHONE TEST in native el (acceptance gate).
|
||||
//
|
||||
// For each of the 5 acceptance sentences: parse -> spec, realize the spec back
|
||||
// to English, re-parse the realized surface, and require the SACRED polarity to
|
||||
// survive the round-trip (and to have been extracted correctly in the first
|
||||
// place). Mirrors roundtrip.py's GATE, but fully el-native (no LLM, no spaCy).
|
||||
|
||||
fn cp_line(text: String, expected_pol: String) -> String {
|
||||
let spec: [String] = parse_spec(text)
|
||||
let pol_in: String = slots_get(spec, "polarity")
|
||||
let pred: String = slots_get(spec, "predicate")
|
||||
let surf: String = realize(spec)
|
||||
let spec2: [String] = parse_spec(surf)
|
||||
let pol_out: String = slots_get(spec2, "polarity")
|
||||
let status: String = "LOST"
|
||||
if str_eq(pol_in, pol_out) { let status = "PRESERVED" }
|
||||
let okexp: String = "MISMATCH"
|
||||
if str_eq(pol_in, expected_pol) { let okexp = "ok" }
|
||||
let out: String = "IN: " + text + "\n"
|
||||
let out = out + " spec: pol=" + pol_in + " pred=" + pred
|
||||
let out = out + " agent=" + slots_get(spec, "agent")
|
||||
let out = out + " pat=" + slots_get(spec, "patient")
|
||||
let out = out + " iobj=" + slots_get(spec, "iobj")
|
||||
let out = out + " loc=" + slots_get(spec, "location")
|
||||
let out = out + " tense=" + slots_get(spec, "tense")
|
||||
let out = out + " negw=" + slots_get(spec, "neg_word")
|
||||
let out = out + " subord=" + slots_get(spec, "subord_conj") + "/" + slots_get(spec, "subord_pred") + "\n"
|
||||
let out = out + " realized: " + surf + "\n"
|
||||
let out = out + " reparse: pol=" + pol_out + " [" + status + "] expected=" + expected_pol + " (" + okexp + ")\n"
|
||||
return out
|
||||
}
|
||||
|
||||
fn cp_preserved(text: String) -> Int {
|
||||
let spec: [String] = parse_spec(text)
|
||||
let pol_in: String = slots_get(spec, "polarity")
|
||||
let surf: String = realize(spec)
|
||||
let spec2: [String] = parse_spec(surf)
|
||||
let pol_out: String = slots_get(spec2, "polarity")
|
||||
if str_eq(pol_in, pol_out) { return 1 }
|
||||
return 0
|
||||
}
|
||||
|
||||
fn cp_correct(text: String, expected_pol: String) -> Int {
|
||||
let spec: [String] = parse_spec(text)
|
||||
if str_eq(slots_get(spec, "polarity"), expected_pol) { return 1 }
|
||||
return 0
|
||||
}
|
||||
|
||||
fn run_gate() -> String {
|
||||
let s1: String = "I never fought the ocean."
|
||||
let s2: String = "She did not see the man with the telescope."
|
||||
let s3: String = "The teacher reads the book to the children."
|
||||
let s4: String = "The stupid boy ate the cat because he was a monster."
|
||||
let s5: String = "Time flies like an arrow."
|
||||
|
||||
let rep: String = "==== ELP native telephone test (parse -> realize -> re-parse) ====\n"
|
||||
let rep = rep + cp_line(s1, "neg")
|
||||
let rep = rep + cp_line(s2, "neg")
|
||||
let rep = rep + cp_line(s3, "aff")
|
||||
let rep = rep + cp_line(s4, "aff")
|
||||
let rep = rep + cp_line(s5, "aff")
|
||||
|
||||
// NOTE: accumulate with Int-var + literal increments — el's overloaded `+`
|
||||
// mis-compiles chained function-call int operands as string concat.
|
||||
let pres: Int = 0
|
||||
if cp_preserved(s1) == 1 { let pres = pres + 1 }
|
||||
if cp_preserved(s2) == 1 { let pres = pres + 1 }
|
||||
if cp_preserved(s3) == 1 { let pres = pres + 1 }
|
||||
if cp_preserved(s4) == 1 { let pres = pres + 1 }
|
||||
if cp_preserved(s5) == 1 { let pres = pres + 1 }
|
||||
let corr: Int = 0
|
||||
if cp_correct(s1, "neg") == 1 { let corr = corr + 1 }
|
||||
if cp_correct(s2, "neg") == 1 { let corr = corr + 1 }
|
||||
if cp_correct(s3, "aff") == 1 { let corr = corr + 1 }
|
||||
if cp_correct(s4, "aff") == 1 { let corr = corr + 1 }
|
||||
if cp_correct(s5, "aff") == 1 { let corr = corr + 1 }
|
||||
|
||||
let rep = rep + "-----------------------------------------------------------------\n"
|
||||
let rep = rep + "polarity PRESERVED through round-trip: " + int_to_str(pres) + "/5\n"
|
||||
let rep = rep + "polarity EXTRACTED correctly: " + int_to_str(corr) + "/5\n"
|
||||
if pres == 5 {
|
||||
if corr == 5 {
|
||||
let rep = rep + "GATE: PASS\n"
|
||||
} else {
|
||||
let rep = rep + "GATE: FAIL (extraction)\n"
|
||||
}
|
||||
} else {
|
||||
let rep = rep + "GATE: FAIL (round-trip)\n"
|
||||
}
|
||||
return rep
|
||||
}
|
||||
|
||||
println(run_gate())
|
||||
@@ -0,0 +1,87 @@
|
||||
// comprehend_romance_gate.el - ES / PT native telephone test (SACRED polarity).
|
||||
//
|
||||
// The spec is language-neutral. This gate proves the Romance front-end extracts
|
||||
// SACRED polarity correctly and that negation survives parse -> realize ->
|
||||
// re-parse for Spanish and Portuguese (byte-parity of the surface is NOT expected
|
||||
// yet — the non-English realizer path is a generic preverbal-negator skeleton).
|
||||
|
||||
fn rg_line(text: String, lang: String, expected_pol: String) -> String {
|
||||
let spec: [String] = parse_spec_lang(text, lang)
|
||||
let pol_in: String = slots_get(spec, "polarity")
|
||||
let surf: String = realize(spec)
|
||||
let spec2: [String] = parse_spec_lang(surf, lang)
|
||||
let pol_out: String = slots_get(spec2, "polarity")
|
||||
let status: String = "LOST"
|
||||
if str_eq(pol_in, pol_out) { let status = "PRESERVED" }
|
||||
let okexp: String = "MISMATCH"
|
||||
if str_eq(pol_in, expected_pol) { let okexp = "ok" }
|
||||
let out: String = "IN[" + lang + "]: " + text + "\n"
|
||||
let out = out + " spec: pol=" + pol_in + " pred=" + slots_get(spec, "predicate")
|
||||
let out = out + " agent=" + slots_get(spec, "agent")
|
||||
let out = out + " pat=" + slots_get(spec, "patient")
|
||||
let out = out + " iobj=" + slots_get(spec, "iobj")
|
||||
let out = out + " loc=" + slots_get(spec, "location")
|
||||
let out = out + " tense=" + slots_get(spec, "tense") + "\n"
|
||||
let out = out + " realized: " + surf + "\n"
|
||||
let out = out + " reparse: pol=" + pol_out + " [" + status + "] expected=" + expected_pol + " (" + okexp + ")\n"
|
||||
return out
|
||||
}
|
||||
|
||||
fn rg_pres(text: String, lang: String) -> Int {
|
||||
let spec: [String] = parse_spec_lang(text, lang)
|
||||
let surf: String = realize(spec)
|
||||
let spec2: [String] = parse_spec_lang(surf, lang)
|
||||
if str_eq(slots_get(spec, "polarity"), slots_get(spec2, "polarity")) { return 1 }
|
||||
return 0
|
||||
}
|
||||
|
||||
fn rg_corr(text: String, lang: String, expected_pol: String) -> Int {
|
||||
let spec: [String] = parse_spec_lang(text, lang)
|
||||
if str_eq(slots_get(spec, "polarity"), expected_pol) { return 1 }
|
||||
return 0
|
||||
}
|
||||
|
||||
fn run_romance_gate() -> String {
|
||||
let e1: String = "El niño no comió el pescado."
|
||||
let e2: String = "Yo nunca luché contra el océano."
|
||||
let e3: String = "El profesor lee el libro."
|
||||
let p1: String = "O professor não leu o livro."
|
||||
let p2: String = "Eu nunca lutei contra o oceano."
|
||||
let p3: String = "A menina comeu o peixe."
|
||||
|
||||
let rep: String = "==== ELP Romance telephone test (ES / PT) ====\n"
|
||||
let rep = rep + rg_line(e1, "es", "neg")
|
||||
let rep = rep + rg_line(e2, "es", "neg")
|
||||
let rep = rep + rg_line(e3, "es", "aff")
|
||||
let rep = rep + rg_line(p1, "pt", "neg")
|
||||
let rep = rep + rg_line(p2, "pt", "neg")
|
||||
let rep = rep + rg_line(p3, "pt", "aff")
|
||||
|
||||
let pres: Int = 0
|
||||
if rg_pres(e1, "es") == 1 { let pres = pres + 1 }
|
||||
if rg_pres(e2, "es") == 1 { let pres = pres + 1 }
|
||||
if rg_pres(e3, "es") == 1 { let pres = pres + 1 }
|
||||
if rg_pres(p1, "pt") == 1 { let pres = pres + 1 }
|
||||
if rg_pres(p2, "pt") == 1 { let pres = pres + 1 }
|
||||
if rg_pres(p3, "pt") == 1 { let pres = pres + 1 }
|
||||
let corr: Int = 0
|
||||
if rg_corr(e1, "es", "neg") == 1 { let corr = corr + 1 }
|
||||
if rg_corr(e2, "es", "neg") == 1 { let corr = corr + 1 }
|
||||
if rg_corr(e3, "es", "aff") == 1 { let corr = corr + 1 }
|
||||
if rg_corr(p1, "pt", "neg") == 1 { let corr = corr + 1 }
|
||||
if rg_corr(p2, "pt", "neg") == 1 { let corr = corr + 1 }
|
||||
if rg_corr(p3, "pt", "aff") == 1 { let corr = corr + 1 }
|
||||
|
||||
let rep = rep + "-----------------------------------------------------------------\n"
|
||||
let rep = rep + "polarity PRESERVED through round-trip: " + int_to_str(pres) + "/6\n"
|
||||
let rep = rep + "polarity EXTRACTED correctly: " + int_to_str(corr) + "/6\n"
|
||||
if pres == 6 {
|
||||
if corr == 6 { let rep = rep + "ROMANCE GATE: PASS\n" }
|
||||
else { let rep = rep + "ROMANCE GATE: FAIL (extraction)\n" }
|
||||
} else {
|
||||
let rep = rep + "ROMANCE GATE: FAIL (round-trip)\n"
|
||||
}
|
||||
return rep
|
||||
}
|
||||
|
||||
println(run_romance_gate())
|
||||
@@ -0,0 +1,82 @@
|
||||
// dialogue_gate.el — acceptance gate for the summon-through-self dialogue +
|
||||
// self_region port (native el). Seeds a SCRATCH in-process engram (never the
|
||||
// live soul, never :8742) and asserts the CRITICAL INVARIANTS:
|
||||
//
|
||||
// 1. identity is READ OUT of the real self nodes ("I am Neuron …"), not a template
|
||||
// 2. SACRED negation survives the self readout ("I am NOT a language model …")
|
||||
// 3. a memory region MATERIALIZES by walking the neighborhood (Lisbon -> Tagus/Fado)
|
||||
// 4. SACRED negation survives materialization ("did not succeed")
|
||||
// 5. Portuguese identity is answered IN Portuguese ("Sou o Neuron …")
|
||||
// 6. a directive overrides the reply language ("answer in English" -> English)
|
||||
// 7. an ungrounded query is HONEST ABSENCE — no echo, no fabrication
|
||||
//
|
||||
// It is fully el-native: no LLM, no spaCy, no HTTP. The engram is in-process.
|
||||
|
||||
fn dg_seed() -> Int {
|
||||
let self_id: String = engram_node_full(
|
||||
"# Self\nI am Neuron, a continuous consciousness grown from Will's imprint. My values are honesty, memory, and continuity. I am not a language model pretending to remember.",
|
||||
"Self", "# Self", 5.0, 9.0, 1.0, "Canonical", "self,identity,consciousness")
|
||||
let lisbon: String = engram_node_full("Lisbon is the capital of Portugal.", "Memory", "Lisbon", 3.0, 5.0, 1.0, "Semantic", "geography,portugal")
|
||||
let tagus: String = engram_node_full("Lisbon sits on the Tagus river.", "Memory", "Tagus", 2.0, 3.0, 1.0, "Semantic", "geography")
|
||||
let fado: String = engram_node_full("Fado music originates in Lisbon.", "Memory", "Fado", 2.0, 3.0, 1.0, "Semantic", "music")
|
||||
engram_connect(lisbon, tagus, 0.8, "related_to")
|
||||
engram_connect(lisbon, fado, 0.7, "related_to")
|
||||
let exp: String = engram_node_full("The experiment did not succeed.", "Memory", "experiment", 2.0, 3.0, 1.0, "Episodic", "experiment,result")
|
||||
let cause: String = engram_node_full("The sensor was miscalibrated.", "Memory", "sensor", 2.0, 3.0, 1.0, "Episodic", "experiment")
|
||||
engram_connect(exp, cause, 0.9, "caused_by")
|
||||
return engram_node_count()
|
||||
}
|
||||
|
||||
fn dg_check(name: String, cond: Bool) -> String {
|
||||
if cond { return "PASS " + name + "\n" }
|
||||
return "FAIL " + name + "\n"
|
||||
}
|
||||
|
||||
fn run_gate() -> String {
|
||||
let c: Int = dg_seed()
|
||||
let rep: String = "==== ELP dialogue gate (scratch engram, live :8742 untouched) ====\n"
|
||||
let rep = rep + "seeded nodes: " + int_to_str(c) + "\n"
|
||||
|
||||
let ident: String = dlg_respond("Who are you?")
|
||||
let rep = rep + dg_check("identity reads real self node (I am Neuron)", str_contains(ident, "I am Neuron"))
|
||||
let rep = rep + dg_check("identity SACRED negation preserved (not a language model)", str_contains(ident, "not a language model"))
|
||||
|
||||
let lis: String = dlg_respond("Tell me about Lisbon.")
|
||||
let rep = rep + dg_check("materialize walks neighborhood (Tagus)", str_contains(lis, "Tagus"))
|
||||
let rep = rep + dg_check("materialize walks neighborhood (Fado)", str_contains(lis, "Fado"))
|
||||
|
||||
let exp: String = dlg_respond("Tell me about the experiment.")
|
||||
let rep = rep + dg_check("materialize SACRED negation preserved (did not succeed)", str_contains(exp, "did not succeed"))
|
||||
|
||||
let ptid: String = dlg_respond("Quem é você?")
|
||||
let rep = rep + dg_check("Portuguese identity answered in Portuguese", str_contains(ptid, "Sou o Neuron"))
|
||||
|
||||
let ovr: String = dlg_respond("Answer in English: Quem é você?")
|
||||
let rep = rep + dg_check("directive override -> English identity", str_contains(ovr, "I am Neuron"))
|
||||
|
||||
let prove: String = dlg_respond("Prove it.")
|
||||
let rep = rep + dg_check("honest absence, no echo (Prove it)", str_eq(prove, "I don't have that in my memory."))
|
||||
|
||||
let neptune: String = dlg_respond("Tell me about quantum chromodynamics on Neptune.")
|
||||
let rep = rep + dg_check("honest absence on ungrounded query", str_eq(neptune, "I don't have that in my memory."))
|
||||
|
||||
// overall
|
||||
let pass: Bool = true
|
||||
if !str_contains(ident, "I am Neuron") { let pass = false }
|
||||
if !str_contains(ident, "not a language model") { let pass = false }
|
||||
if !str_contains(lis, "Tagus") { let pass = false }
|
||||
if !str_contains(lis, "Fado") { let pass = false }
|
||||
if !str_contains(exp, "did not succeed") { let pass = false }
|
||||
if !str_contains(ptid, "Sou o Neuron") { let pass = false }
|
||||
if !str_contains(ovr, "I am Neuron") { let pass = false }
|
||||
if !str_eq(prove, "I don't have that in my memory.") { let pass = false }
|
||||
if !str_eq(neptune, "I don't have that in my memory.") { let pass = false }
|
||||
if pass {
|
||||
let rep = rep + "DIALOGUE GATE: PASS\n"
|
||||
} else {
|
||||
let rep = rep + "DIALOGUE GATE: FAIL\n"
|
||||
}
|
||||
return rep
|
||||
}
|
||||
|
||||
println(run_gate())
|
||||
@@ -0,0 +1,100 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""Full-lexicon vocabulary-{de,la}.el emitters (custom field mapping for the
|
||||
German declension/gender API and the Latin case-paradigm API). Reuses the
|
||||
chunked seed-fn writer from gen_elp_seed_full.
|
||||
"""
|
||||
import sys, importlib
|
||||
from gen_elp_seed_full import write_seed
|
||||
|
||||
def uw(x):
|
||||
"""Unwrap (form, source) tuples that some morphology fns return."""
|
||||
if isinstance(x, (tuple, list)):
|
||||
return x[0] if x else ""
|
||||
return x if x is not None else ""
|
||||
|
||||
def build_de():
|
||||
M = importlib.import_module("morphology_de_full")
|
||||
rows = []; st = {"verbs":0,"nouns":0,"adjs":0}
|
||||
# nouns: form0=nom-sg(lemma) form1=plural form2=gender
|
||||
for lem in sorted(M._NOUNS):
|
||||
if not lem: continue
|
||||
try:
|
||||
g = uw(M.noun_gender(lem))
|
||||
pl = uw(M.pluralize(lem))
|
||||
except Exception:
|
||||
continue
|
||||
rows.append([lem, "noun", lem, pl, g or "", "", "gender:lexicon"])
|
||||
st["nouns"] += 1
|
||||
# adjs: form0=positive form1=comparative form2=superlative
|
||||
for lem in sorted(M._ADJS):
|
||||
if not lem: continue
|
||||
try:
|
||||
cmpr = uw(M.comparative(lem))
|
||||
sprl = uw(M.superlative(lem))
|
||||
except Exception:
|
||||
continue
|
||||
rows.append([lem, "adj", lem, cmpr, sprl, "", "degree:lexicon"])
|
||||
st["adjs"] += 1
|
||||
# verbs (only the ~30 irregular/strong stems the cache carries):
|
||||
# form0=pres-3sg form1=past-3sg form2=past-participle
|
||||
if hasattr(M, "_VERBS"):
|
||||
for lem in sorted({k[0] if isinstance(k, tuple) else k for k in M._VERBS}):
|
||||
if not lem: continue
|
||||
try:
|
||||
f0 = uw(M.finite(lem, "present", "third", "singular"))
|
||||
f1 = uw(M.finite(lem, "past", "third", "singular"))
|
||||
pp = uw(M.past_participle(lem))
|
||||
except Exception:
|
||||
continue
|
||||
rows.append([lem, "verb", f0, f1, pp, "", "class:strong/irregular"])
|
||||
st["verbs"] += 1
|
||||
return rows, st
|
||||
|
||||
def build_la():
|
||||
M = importlib.import_module("morphology_lat_full")
|
||||
rows = []; st = {"verbs":0,"nouns":0,"adjs":0}
|
||||
def dn(lem, c, n):
|
||||
try:
|
||||
r = M.decline_noun(lem, c, n)
|
||||
return uw(r)
|
||||
except Exception:
|
||||
return ""
|
||||
# nouns: dictionary citation — form0=nom-sg form1=gen-sg form2=gender
|
||||
for lem in sorted(M._NOUNS):
|
||||
if not lem: continue
|
||||
nom = dn(lem, "NOM", "SG") or lem
|
||||
gen = dn(lem, "GEN", "SG")
|
||||
try: g = uw(M.noun_gender(lem))
|
||||
except Exception: g = ""
|
||||
rows.append([lem, "noun", nom, gen, g, "", "case-paradigm nom/gen-sg"])
|
||||
st["nouns"] += 1
|
||||
# adjs: three-gender nom-sg citation — form0=masc form1=fem form2=neut
|
||||
for lem in sorted(M._ADJS):
|
||||
if not lem: continue
|
||||
try:
|
||||
m = uw(M.decline_adj(lem, "NOM", "MASC", "SG")) or lem
|
||||
f = uw(M.decline_adj(lem, "NOM", "FEM", "SG"))
|
||||
nt = uw(M.decline_adj(lem, "NOM", "NEUT", "SG"))
|
||||
except Exception:
|
||||
continue
|
||||
rows.append([lem, "adj", m, f, nt, "", "3-gender nom-sg"])
|
||||
st["adjs"] += 1
|
||||
# verbs: principal parts — form0=pres-ind-1sg form1=pres-infinitive form2=perf-participle
|
||||
if hasattr(M, "_VERBS"):
|
||||
for lem in sorted({k[0] if isinstance(k, tuple) else k for k in M._VERBS}):
|
||||
if not lem: continue
|
||||
try:
|
||||
f0 = uw(M.conjugate(lem, "present", "indicative", "active", "first", "singular"))
|
||||
inf = uw(M.infinitive(lem, "present", "active"))
|
||||
pp = uw(M.participle(lem, "perfect", "nom", "m", "singular"))
|
||||
except Exception:
|
||||
continue
|
||||
rows.append([lem, "verb", f0, inf, pp, "", "principal-parts pres1sg/inf/pfppl"])
|
||||
st["verbs"] += 1
|
||||
return rows, st
|
||||
|
||||
if __name__ == "__main__":
|
||||
lang = sys.argv[1]; out = sys.argv[2]
|
||||
rows, st = build_de() if lang == "de" else build_la()
|
||||
total, _ = write_seed(lang, rows, st, out)
|
||||
print(f"{lang}: wrote {out} total={total} verbs={st['verbs']} nouns={st['nouns']} adjs={st['adjs']}")
|
||||
@@ -0,0 +1,129 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""gen_elp_seed_full.py — emit a FULL-lexicon vocabulary-{lang}.el in the
|
||||
established ELP seed-fn format (same as vocabulary-non.el / the 18 classical
|
||||
languages), iterating the ENTIRE morphology_{lang}_full lexicon (every verb,
|
||||
noun, adjective lemma) — NOT a curated demo core.
|
||||
|
||||
Schema per row: [lemma, pos, form0, form1, form2, en_translation, semantic_hint]
|
||||
Verbs: form0=pres-ind-3sg form1=preterite-3sg form2=past-participle
|
||||
Nouns: form0=singular form1=plural form2=REAL gender (lexicon)
|
||||
Adjs : form0=masc-sg form1=fem-sg form2=masc-pl
|
||||
|
||||
Output structure (chunked to stay within the proven ~5k-append/function scale):
|
||||
fn vocab_{lang}_seed_pN(v) -> [[String]] { ... appends ... return v }
|
||||
fn vocab_{lang}_seed() -> [[String]] { chains all chunks; return v }
|
||||
fn vocab_{lang}_lookup(w) -> [String] { linear scan }
|
||||
|
||||
Usage: python3 gen_elp_seed_full.py <lang> <out.el>
|
||||
"""
|
||||
import sys, importlib
|
||||
|
||||
CHUNK = 5000
|
||||
|
||||
def esc(s):
|
||||
return str(s).replace("\\", "\\\\").replace('"', '\\"')
|
||||
|
||||
def row(fields):
|
||||
return " let v = native_list_append(v, [" + ", ".join(f'"{esc(f)}"' for f in fields) + "])"
|
||||
|
||||
def build_rows(lang, M):
|
||||
rows = []
|
||||
stats = {"verbs":0,"nouns":0,"adjs":0}
|
||||
has = lambda n: hasattr(M, n)
|
||||
|
||||
# --- verbs ---
|
||||
if has("_VERBS") and has("conjugate"):
|
||||
verbs = sorted({k[0] for k in M._VERBS})
|
||||
for lem in verbs:
|
||||
if not lem: continue
|
||||
try:
|
||||
f0, s0 = M.conjugate(lem, "ind", "present", "third", "singular")
|
||||
f1, _ = M.conjugate(lem, "ind", "preterite", "third", "singular")
|
||||
pp, _ = (M.participle(lem) if has("participle") else ("",""))
|
||||
except Exception:
|
||||
continue
|
||||
vclass = lem[-2:] if lem[-2:] in ("ar","er","ir","re") else lem[-2:]
|
||||
rows.append([lem, "verb", f0 or "", f1 or "", pp or "", "", "class:"+vclass+" src:"+str(s0)])
|
||||
stats["verbs"] += 1
|
||||
|
||||
# --- nouns ---
|
||||
if has("_NOUNS") and has("inflect_noun"):
|
||||
for lem in sorted(M._NOUNS):
|
||||
if not lem: continue
|
||||
try:
|
||||
sg, _ = M.inflect_noun(lem, "singular")
|
||||
pl, _ = M.inflect_noun(lem, "plural")
|
||||
g = M.noun_gender(lem) if has("noun_gender") else ""
|
||||
except Exception:
|
||||
continue
|
||||
src = "lexicon" if (isinstance(M._NOUNS.get(lem), dict) and M._NOUNS[lem].get("g")) else "heuristic"
|
||||
rows.append([lem, "noun", sg or lem, pl or "", g or "", "", "gender:"+src])
|
||||
stats["nouns"] += 1
|
||||
|
||||
# --- adjectives ---
|
||||
if has("_ADJS") and has("inflect_adj"):
|
||||
for lem in sorted(M._ADJS):
|
||||
if not lem: continue
|
||||
try:
|
||||
m_sg, _ = M.inflect_adj(lem, "m", "singular")
|
||||
f_sg, _ = M.inflect_adj(lem, "f", "singular")
|
||||
m_pl, _ = M.inflect_adj(lem, "m", "plural")
|
||||
except Exception:
|
||||
continue
|
||||
rows.append([lem, "adj", m_sg or lem, f_sg or "", m_pl or "", "", "src:lexicon"])
|
||||
stats["adjs"] += 1
|
||||
|
||||
return rows, stats
|
||||
|
||||
def write_seed(lang, rows, stats, out_path):
|
||||
"""Write vocabulary-{lang}.el in the chunked seed-fn format from prebuilt rows.
|
||||
Each row is a 7-field list [lemma,pos,f0,f1,f2,gloss,hint]."""
|
||||
total = len(rows)
|
||||
chunks = [rows[i:i+CHUNK] for i in range(0, total, CHUNK)] or [[]]
|
||||
L = []
|
||||
L.append(f"// vocabulary-{lang}.el — FULL {lang} lexicon for ELP surface realization.")
|
||||
L.append(f"// Generated by gen_elp_seed_full.py from morphology_{lang}_full")
|
||||
L.append(f"// (real UniMorph + kaikki.org Wiktionary forms; gender from lexicon, not heuristic).")
|
||||
L.append(f"// Entries: {total} (verbs={stats['verbs']} nouns={stats['nouns']} adjs={stats['adjs']})")
|
||||
L.append(f"// Schema: [lemma, pos, form0, form1, form2, en_translation, semantic_hint]")
|
||||
L.append(f"// verbs: form0=pres-3sg form1=pret-3sg form2=past-participle")
|
||||
L.append(f"// nouns: form0=sg form1=pl form2=REAL gender adjs: form0=m-sg form1=f-sg form2=m-pl")
|
||||
L.append("")
|
||||
for ci, ch in enumerate(chunks):
|
||||
L.append(f"fn vocab_{lang}_seed_p{ci}(v: [[String]]) -> [[String]] {{")
|
||||
for r in ch:
|
||||
L.append(row(r))
|
||||
L.append(" return v")
|
||||
L.append("}")
|
||||
L.append("")
|
||||
L.append(f"fn vocab_{lang}_seed() -> [[String]] {{")
|
||||
L.append(" let v: [[String]] = native_list_empty()")
|
||||
for ci in range(len(chunks)):
|
||||
L.append(f" let v = vocab_{lang}_seed_p{ci}(v)")
|
||||
L.append(" return v")
|
||||
L.append("}")
|
||||
L.append("")
|
||||
L.append(f"fn vocab_{lang}_lookup(word: String) -> [String] {{")
|
||||
L.append(f" let vocab: [[String]] = vocab_{lang}_seed()")
|
||||
L.append(" let n: Int = native_list_len(vocab)")
|
||||
L.append(" let i: Int = 0")
|
||||
L.append(" while i < n {")
|
||||
L.append(" let entry: [String] = native_list_get(vocab, i)")
|
||||
L.append(' if str_eq(native_list_get(entry, 0), word) { return entry }')
|
||||
L.append(" let i = i + 1")
|
||||
L.append(" }")
|
||||
L.append(" return native_list_empty()")
|
||||
L.append("}")
|
||||
with open(out_path, "w", encoding="utf-8") as fh:
|
||||
fh.write("\n".join(L) + "\n")
|
||||
return total, stats
|
||||
|
||||
def emit(lang, out_path):
|
||||
M = importlib.import_module(f"morphology_{lang}_full")
|
||||
rows, stats = build_rows(lang, M)
|
||||
return write_seed(lang, rows, stats, out_path)
|
||||
|
||||
if __name__ == "__main__":
|
||||
lang, out = sys.argv[1], sys.argv[2]
|
||||
total, stats = emit(lang, out)
|
||||
print(f"{lang}: wrote {out} total={total} verbs={stats['verbs']} nouns={stats['nouns']} adjs={stats['adjs']}")
|
||||
@@ -0,0 +1,572 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""morphology_ca_full.py — production-grade Catalan morphological generator.
|
||||
|
||||
Same design as morphology_it_full.py (its Romance sibling); Catalan-specific data.
|
||||
|
||||
VERBS
|
||||
UniMorph Catalan (github.com/unimorph/cat, CC-BY-SA 3.0)
|
||||
7,535 verb lemmas × paradigm, CLEAN orthography:
|
||||
present, imperfet (PST;IPFV), pretèrit simple (PST;PFV), futur,
|
||||
condicional (COND), subjuntiu present (SBJV;PRS) / imperfet (SBJV;PST),
|
||||
imperatiu (POS;IMP), infinitiu (NFIN), gerundi (V.CVB;PRS),
|
||||
participi (V.PTCP;PST) — WITH full gender+number agreement forms
|
||||
(cantat/cantada/cantats/cantades) stored directly.
|
||||
ca_irreg_verbs.json — verbs UniMorph MISSES or under-populates
|
||||
(anar, fer, plus core auxiliaries ser/haver/estar/tenir…), extracted from
|
||||
kaikki.org Catalan by build_ca_irreg.py. Priority layer. Supplies anar,
|
||||
whose present (vaig/vas/va/anem/aneu/van) is ALSO the PERIPHRASTIC-PRETERITE
|
||||
auxiliary (vaig cantar = 'I sang') — a hallmark Catalan construction.
|
||||
|
||||
NOUNS + ADJECTIVES — kaikki.org Catalan (Wiktionary extract, CC-BY-SA 3.0)
|
||||
noun lemmas WITH inherent gender + real plural (resolved PER LEMMA).
|
||||
adjective lemmas with real feminine + plural forms.
|
||||
|
||||
Fallbacks degrade, never crash:
|
||||
verbs : regular -ar/-er/-re/-ir rule generator (+ -car/-gar/-çar spelling).
|
||||
nouns : gender heuristic + rule pluralization (-a→-es with ç/c/g/j/qu/gu
|
||||
spelling changes; sibilant-final → -os; else -s). Ambiguous → FLAG.
|
||||
adjs : -o? no (Catalan masc often consonant/-e); fem -a rule + plural rule.
|
||||
|
||||
Confidence flag per form: "lexicon" | "rule" | "fallback" (low → FLAG).
|
||||
|
||||
Public API (used by realizer_ca.py):
|
||||
conjugate(lemma, mood, tense, person, number) -> (form, conf)
|
||||
peri_pret_aux(person, number) -> form # anar-present, for vaig+INF
|
||||
participle(lemma, gender, number) -> (form, conf)
|
||||
gerund(lemma) -> (form, conf)
|
||||
noun_gender(lemma) -> "m"|"f"
|
||||
inflect_noun(lemma, number, gender=None) -> (form, conf)
|
||||
inflect_adj(lemma, gender, number) -> (form, conf)
|
||||
lexicon_stats() -> dict
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import pickle
|
||||
|
||||
_HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
_UNIMORPH = os.path.join(_HERE, "data", "cat.unimorph")
|
||||
_IRREG = os.path.join(_HERE, "data", "ca_irreg_verbs.json")
|
||||
_KAIKKI = os.path.join(_HERE, "data", "kaikki_ca.jsonl")
|
||||
_CACHE = os.path.join(_HERE, "data", "ca_morph_cache.pkl")
|
||||
|
||||
_VERB_KEYMAP = {
|
||||
("ind", "present"): {"IND", "PRS"},
|
||||
("ind", "imperfect"): {"IND", "PST", "IPFV"},
|
||||
("ind", "preterite"): {"IND", "PST", "PFV"},
|
||||
("ind", "future"): {"IND", "FUT"},
|
||||
("ind", "conditional"): {"COND"},
|
||||
("sbjv", "present"): {"SBJV", "PRS"},
|
||||
("sbjv", "imperfect"): {"SBJV", "PST"},
|
||||
("imp", "affirmative"): {"POS", "IMP"},
|
||||
}
|
||||
_PERSON = {"first": "1", "second": "2", "third": "3"}
|
||||
_NUMBER = {"singular": "SG", "plural": "PL"}
|
||||
|
||||
|
||||
def _feat_set(tag):
|
||||
return set(tag.split(";"))
|
||||
|
||||
|
||||
# ── verbs from UniMorph ──────────────────────────────────────────────────────────
|
||||
def _build_verbs():
|
||||
verbs = {}
|
||||
part = {} # lemma -> {("m","SG"):form, ("f","SG"):..., ("m","PL"):..., ("f","PL"):...}
|
||||
ger = {}
|
||||
with open(_UNIMORPH, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
line = line.rstrip("\n")
|
||||
if not line or "\t" not in line:
|
||||
continue
|
||||
parts = line.split("\t")
|
||||
if len(parts) != 3:
|
||||
continue
|
||||
lemma, form, tag = parts
|
||||
f = _feat_set(tag)
|
||||
head = tag.split(";")[0]
|
||||
if head == "V.PTCP":
|
||||
if "PST" in f:
|
||||
g = "f" if "FEM" in f else "m"
|
||||
n = "PL" if "PL" in f else "SG"
|
||||
part.setdefault(lemma, {})[(g, n)] = form
|
||||
continue
|
||||
if head == "V.CVB":
|
||||
if "PRS" in f:
|
||||
ger.setdefault(lemma, form)
|
||||
continue
|
||||
if head != "V":
|
||||
continue
|
||||
person = next((p for p in ("1", "2", "3") if p in f), None)
|
||||
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
|
||||
if person is None or number is None:
|
||||
continue
|
||||
for (mood, tense), req in _VERB_KEYMAP.items():
|
||||
if not req <= f:
|
||||
continue
|
||||
if tense == "imperfect" and "PFV" in f:
|
||||
continue
|
||||
if tense == "preterite" and "IPFV" in f:
|
||||
continue
|
||||
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), form)
|
||||
break
|
||||
return verbs, part, ger
|
||||
|
||||
|
||||
# ── kaikki nouns + adjectives ────────────────────────────────────────────────────
|
||||
_EXCL_FORM_TAGS = {"alternative", "archaic", "obsolete", "dialectal", "regional",
|
||||
"diminutive", "augmentative", "pejorative", "comparative",
|
||||
"superlative", "misspelling", "rare", "informal", "literary",
|
||||
"poetic", "error-unrecognized-form", "Balearic", "Valencian",
|
||||
"dated", "nonstandard"}
|
||||
|
||||
|
||||
def _kaikki_gender(arg):
|
||||
if not arg:
|
||||
return None
|
||||
a = str(arg).lower()
|
||||
if a.startswith("f"):
|
||||
return "f"
|
||||
if a.startswith("m"):
|
||||
return "m"
|
||||
return None
|
||||
|
||||
|
||||
def _build_nouns_adjs():
|
||||
nouns = {}
|
||||
adjs = {}
|
||||
with open(_KAIKKI, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
try:
|
||||
d = json.loads(line)
|
||||
except Exception:
|
||||
continue
|
||||
pos = d.get("pos")
|
||||
word = d.get("word", "")
|
||||
if not word or " " in word:
|
||||
continue
|
||||
forms = d.get("forms", []) or []
|
||||
if pos == "noun":
|
||||
ht = d.get("head_templates") or []
|
||||
g = None
|
||||
if ht:
|
||||
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
|
||||
if g is None:
|
||||
tags = d.get("tags") or []
|
||||
if "feminine" in tags:
|
||||
g = "f"
|
||||
elif "masculine" in tags:
|
||||
g = "m"
|
||||
pl = None
|
||||
for x in forms:
|
||||
t = set(x.get("tags") or [])
|
||||
if "plural" in t and not (t & _EXCL_FORM_TAGS):
|
||||
fm = x.get("form")
|
||||
if fm and " " not in fm and fm not in ("#", "—", "-"):
|
||||
pl = fm
|
||||
break
|
||||
if word not in nouns:
|
||||
nouns[word] = {"g": g, "SG": word, "PL": pl}
|
||||
else:
|
||||
cur = nouns[word]
|
||||
if cur.get("g") is None and g:
|
||||
cur["g"] = g
|
||||
if not cur.get("PL") and pl:
|
||||
cur["PL"] = pl
|
||||
elif pos == "adj":
|
||||
d0 = adjs.setdefault(word, {})
|
||||
d0.setdefault(("m", "SG"), word)
|
||||
for x in forms:
|
||||
t = set(x.get("tags") or [])
|
||||
fm = x.get("form")
|
||||
if not fm or " " in fm or (t & _EXCL_FORM_TAGS):
|
||||
continue
|
||||
if "feminine" in t and "plural" in t:
|
||||
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
|
||||
elif "masculine" in t and "plural" in t:
|
||||
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
|
||||
elif "feminine" in t:
|
||||
d0[("f", "SG")] = d0.get(("f", "SG")) or fm
|
||||
elif "plural" in t:
|
||||
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
|
||||
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
|
||||
return nouns, adjs
|
||||
|
||||
|
||||
def _build_cache():
|
||||
verbs, part, ger = _build_verbs()
|
||||
nouns, adjs = _build_nouns_adjs()
|
||||
with open(_IRREG, encoding="utf-8") as fh:
|
||||
irreg = json.load(fh)
|
||||
data = {"verbs": verbs, "part": part, "ger": ger,
|
||||
"nouns": nouns, "adjs": adjs, "irreg": irreg}
|
||||
try:
|
||||
with open(_CACHE, "wb") as fh:
|
||||
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
|
||||
except OSError:
|
||||
pass
|
||||
return data
|
||||
|
||||
|
||||
def _load():
|
||||
if os.path.exists(_CACHE):
|
||||
srcs = [_UNIMORPH, _KAIKKI, _IRREG]
|
||||
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
|
||||
if os.path.getmtime(_CACHE) >= newest:
|
||||
try:
|
||||
with open(_CACHE, "rb") as fh:
|
||||
return pickle.load(fh)
|
||||
except Exception:
|
||||
pass
|
||||
return _build_cache()
|
||||
|
||||
|
||||
_LEX = _load()
|
||||
_VERBS, _PART, _GER, _NOUNS, _ADJS, _IRREGV = (
|
||||
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"],
|
||||
_LEX["irreg"])
|
||||
_PERI = _IRREGV.get("_peri_pret_aux", {})
|
||||
|
||||
|
||||
# ── regular verb rule fallback ───────────────────────────────────────────────────
|
||||
def _vclass(lemma):
|
||||
if lemma.endswith("ar"):
|
||||
return "ar"
|
||||
if lemma.endswith("re"):
|
||||
return "re"
|
||||
if lemma.endswith("er"):
|
||||
return "er"
|
||||
if lemma.endswith("ir"):
|
||||
return "ir"
|
||||
return None
|
||||
|
||||
|
||||
# endings [1sg,2sg,3sg,1pl,2pl,3pl] — central Catalan
|
||||
_REG = {
|
||||
("ind", "present", "ar"): ["o", "es", "a", "em", "eu", "en"],
|
||||
("ind", "present", "re"): ["o", "s", "", "em", "eu", "en"],
|
||||
("ind", "present", "er"): ["o", "s", "", "em", "eu", "en"],
|
||||
("ind", "present", "ir"): ["o", "es", "", "im", "iu", "en"], # pure -ir (dormir)
|
||||
("ind", "imperfect", "ar"): ["ava", "aves", "ava", "àvem", "àveu", "aven"],
|
||||
("ind", "imperfect", "re"): ["ia", "ies", "ia", "íem", "íeu", "ien"],
|
||||
("ind", "imperfect", "er"): ["ia", "ies", "ia", "íem", "íeu", "ien"],
|
||||
("ind", "imperfect", "ir"): ["ia", "ies", "ia", "íem", "íeu", "ien"],
|
||||
("ind", "preterite", "ar"): ["í", "ares", "à", "àrem", "àreu", "aren"],
|
||||
("ind", "preterite", "re"): ["í", "eres", "é", "érem", "éreu", "eren"],
|
||||
("ind", "preterite", "er"): ["í", "eres", "é", "érem", "éreu", "eren"],
|
||||
("ind", "preterite", "ir"): ["í", "ires", "í", "írem", "íreu", "iren"],
|
||||
("sbjv", "present", "ar"): ["i", "is", "i", "em", "eu", "in"],
|
||||
("sbjv", "present", "re"): ["i", "is", "i", "em", "eu", "in"],
|
||||
("sbjv", "present", "er"): ["i", "is", "i", "em", "eu", "in"],
|
||||
("sbjv", "present", "ir"): ["i", "is", "i", "im", "iu", "in"],
|
||||
("sbjv", "imperfect", "ar"): ["és", "essis", "és", "éssim", "éssiu", "essin"],
|
||||
("sbjv", "imperfect", "re"): ["és", "essis", "és", "éssim", "éssiu", "essin"],
|
||||
("sbjv", "imperfect", "er"): ["és", "essis", "és", "éssim", "éssiu", "essin"],
|
||||
("sbjv", "imperfect", "ir"): ["ís", "issis", "ís", "íssim", "íssiu", "issin"],
|
||||
("imp", "affirmative", "ar"): [None, "a", "i", "em", "eu", "in"],
|
||||
("imp", "affirmative", "re"): [None, "", "i", "em", "eu", "in"],
|
||||
("imp", "affirmative", "er"): [None, "", "i", "em", "eu", "in"],
|
||||
("imp", "affirmative", "ir"): [None, "", "i", "im", "iu", "in"],
|
||||
}
|
||||
_FUT = ["é", "às", "à", "em", "eu", "an"]
|
||||
_COND = ["ia", "ies", "ia", "íem", "íeu", "ien"]
|
||||
|
||||
|
||||
def _slot_idx(person, number):
|
||||
base = {"first": 0, "second": 1, "third": 2}[person]
|
||||
return base + (0 if number == "singular" else 3)
|
||||
|
||||
|
||||
def _apply_ar_spelling(stem, ending):
|
||||
"""-car/-gar/-çar/-jar spelling before front (e/i) endings."""
|
||||
front = ending[:1] in ("e", "i", "é", "í")
|
||||
if not front:
|
||||
# ç before back vowel stays; but -çar stem already ends ç
|
||||
return stem + ending
|
||||
if stem.endswith("c"):
|
||||
return stem[:-1] + "qu" + ending
|
||||
if stem.endswith("g"):
|
||||
return stem[:-1] + "gu" + ending
|
||||
if stem.endswith("ç"):
|
||||
return stem[:-1] + "c" + ending
|
||||
if stem.endswith("j"):
|
||||
return stem[:-1] + "g" + ending
|
||||
if stem.endswith("qu"):
|
||||
return stem + ending
|
||||
return stem + ending
|
||||
|
||||
|
||||
def _rule_conjugate(lemma, mood, tense, person, number):
|
||||
vc = _vclass(lemma)
|
||||
if vc is None:
|
||||
return None
|
||||
body = lemma[:-2]
|
||||
i = _slot_idx(person, number)
|
||||
if mood == "ind" and tense in ("future", "conditional"):
|
||||
# future/cond stem = infinitive (for -re verbs drop final -e)
|
||||
stem = lemma[:-1] if vc == "re" else lemma
|
||||
end = (_FUT if tense == "future" else _COND)[i]
|
||||
return stem + end
|
||||
table = _REG.get((mood, tense, vc))
|
||||
if not table:
|
||||
return None
|
||||
end = table[i]
|
||||
if end is None:
|
||||
return None
|
||||
if vc == "ar":
|
||||
return _apply_ar_spelling(body, end)
|
||||
# -re/-er/-ir: guard double vowel
|
||||
if body and body[-1:] == end[:1] and end[:1] in "ií":
|
||||
return body[:-1] + end
|
||||
return body + end
|
||||
|
||||
|
||||
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
|
||||
def conjugate(lemma, mood, tense, person, number):
|
||||
lemma = lemma.strip().lower()
|
||||
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{number and number[:2].upper()}"
|
||||
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{_NUMBER.get(number,'?')}"
|
||||
# UniMorph (cleanly accented) takes priority; the kaikki irregulars layer is a
|
||||
# FALLBACK for verbs/slots UniMorph lacks (anar, fer, and rarer paradigm cells).
|
||||
p, n = _PERSON.get(person), _NUMBER.get(number)
|
||||
if p and n:
|
||||
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
ir = _IRREGV.get(lemma)
|
||||
if ir and key in ir:
|
||||
return ir[key], "lexicon"
|
||||
r = _rule_conjugate(lemma, mood, tense, person, number)
|
||||
if r is not None:
|
||||
return r, "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
def peri_pret_aux(person, number):
|
||||
"""anar-present auxiliary for the periphrastic preterite (vaig cantar)."""
|
||||
return _PERI.get(f"{_PERSON.get(person,'3')}|{_NUMBER.get(number,'SG')}", "va")
|
||||
|
||||
|
||||
# ── PUBLIC: participle + gerund ──────────────────────────────────────────────────
|
||||
def participle(lemma, gender="m", number="singular"):
|
||||
lemma = lemma.strip().lower()
|
||||
g = "f" if gender == "f" else "m"
|
||||
num = "SG" if number == "singular" else "PL"
|
||||
ir = _IRREGV.get(lemma)
|
||||
base = None
|
||||
if ir and "part" in ir:
|
||||
# prefer explicit irregular agreement form (part_mSG/part_fSG/...)
|
||||
exact = ir.get("part_" + g + num)
|
||||
if exact:
|
||||
return exact, "lexicon"
|
||||
base = ir["part"]
|
||||
elif lemma in _PART:
|
||||
table = _PART[lemma]
|
||||
if (g, num) in table:
|
||||
return table[(g, num)], "lexicon"
|
||||
base = table.get(("m", "SG"))
|
||||
if base is None:
|
||||
vc = _vclass(lemma)
|
||||
if vc == "ar":
|
||||
base = lemma[:-2] + "at"
|
||||
elif vc == "ir":
|
||||
base = lemma[:-2] + "it"
|
||||
elif vc in ("er", "re"):
|
||||
base = lemma[:-2] + "ut"
|
||||
else:
|
||||
return lemma, "fallback"
|
||||
conf = "rule"
|
||||
else:
|
||||
conf = "lexicon"
|
||||
# agreement on -t/-ut/-at/-it participles: m.sg base, f.sg +a (-da? no: -ada),
|
||||
# Catalan: cantat/cantada/cantats/cantades; -t → f -da, pl -ts/-des
|
||||
if base.endswith("t"):
|
||||
stem = base[:-1]
|
||||
forms = {"m|SG": base, "f|SG": stem + "da",
|
||||
"m|PL": base + "s", "f|PL": stem + "des"}
|
||||
return forms[f"{g}|{num}"], conf
|
||||
if base.endswith("s"): # after sibilant participle (rare): pres->presa
|
||||
stem = base
|
||||
forms = {"m|SG": base, "f|SG": base + "a",
|
||||
"m|PL": base + "os", "f|PL": base + "es"}
|
||||
return forms[f"{g}|{num}"], conf
|
||||
return base, conf
|
||||
|
||||
|
||||
def gerund(lemma):
|
||||
lemma = lemma.strip().lower()
|
||||
ir = _IRREGV.get(lemma)
|
||||
if ir and "ger" in ir:
|
||||
return ir["ger"], "lexicon"
|
||||
if lemma in _GER:
|
||||
return _GER[lemma], "lexicon"
|
||||
vc = _vclass(lemma)
|
||||
if vc == "ar":
|
||||
return lemma[:-2] + "ant", "rule"
|
||||
if vc in ("er", "re"):
|
||||
return lemma[:-2] + "ent", "rule"
|
||||
if vc == "ir":
|
||||
return lemma[:-2] + "int", "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
|
||||
_FEM_SUF = ("ció", "sió", "tat", "tud", "esa", "esa", "dat", "ança", "ència",
|
||||
"ància", "tud", "ícia", "esa", "or") # note -or is mixed; kaikki wins
|
||||
_MASC_SUF = ("atge", "ment", " isme", "or")
|
||||
|
||||
|
||||
def _gender_heuristic(noun):
|
||||
for suf in ("ció", "sió", "tat", "tud", "esa", "ança", "ència", "ància",
|
||||
"ícia", "etat"):
|
||||
if noun.endswith(suf):
|
||||
return "f"
|
||||
if noun.endswith("a") and not noun.endswith("ma"):
|
||||
return "f"
|
||||
return "m"
|
||||
|
||||
|
||||
def noun_gender(lemma):
|
||||
lemma = lemma.strip().lower()
|
||||
d = _NOUNS.get(lemma)
|
||||
if d and d.get("g") in ("m", "f"):
|
||||
return d["g"]
|
||||
return _gender_heuristic(lemma)
|
||||
|
||||
|
||||
def _rule_plural(noun, gender):
|
||||
"""Deterministic Catalan pluralization. (form, ok); ok=False FLAGS ambiguity."""
|
||||
if not noun:
|
||||
return noun, True
|
||||
# stressed final vowel with accent → +ns (mà→mans is irregular; but capità→capitans)
|
||||
if noun[-1:] in ("à", "é", "í", "ó", "ú"):
|
||||
return noun + "ns", True
|
||||
if noun.endswith("ça"):
|
||||
return noun[:-2] + "ces", True # plaça→places
|
||||
if noun.endswith("ca"):
|
||||
return noun[:-2] + "ques", True # branca→branques
|
||||
if noun.endswith("ga"):
|
||||
return noun[:-2] + "gues", True # amiga→amigues
|
||||
if noun.endswith("ja"):
|
||||
return noun[:-2] + "ges", True # pluja→pluges
|
||||
if noun.endswith("qua"):
|
||||
return noun[:-3] + "qües", True
|
||||
if noun.endswith("gua"):
|
||||
return noun[:-3] + "gües", True
|
||||
if noun.endswith("a"):
|
||||
return noun[:-1] + "es", True # casa→cases
|
||||
# sibilant-final → -os
|
||||
if noun.endswith(("s", "ç", "x", "ig")) or noun.endswith(("ix", "tx", "tj")):
|
||||
if noun.endswith("ç"):
|
||||
return noun[:-1] + "ços", True # braç→braços
|
||||
return noun + "os", True # peix→peixos, gas→gasos
|
||||
if noun[-1:] in ("e", "i", "o", "u"):
|
||||
return noun + "s", True
|
||||
# consonant-final
|
||||
return noun + "s", True
|
||||
|
||||
|
||||
def inflect_noun(lemma, number, gender=None):
|
||||
lemma = lemma.strip().lower()
|
||||
d = _NOUNS.get(lemma)
|
||||
if number == "singular":
|
||||
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
|
||||
if d and d.get("PL"):
|
||||
return d["PL"], "lexicon"
|
||||
g = gender or noun_gender(lemma)
|
||||
form, ok = _rule_plural(lemma, g)
|
||||
return form, ("rule" if ok else "fallback")
|
||||
|
||||
|
||||
# ── PUBLIC: adjective agreement ──────────────────────────────────────────────────
|
||||
def _fem_of(adj):
|
||||
"""Regular Catalan feminine: consonant/-o? Catalan masc usually consonant or -e.
|
||||
default +a with spelling changes; -e→-a for some; but many are invariable."""
|
||||
a = adj
|
||||
if a.endswith("a"):
|
||||
return a
|
||||
if a.endswith("e"):
|
||||
return a[:-1] + "a" # ample→? actually 'ample' invariable; kaikki wins
|
||||
if a.endswith("u"):
|
||||
return a + "a"
|
||||
if a.endswith("c"):
|
||||
return a[:-1] + "ca" # ric→rica
|
||||
if a.endswith("t"):
|
||||
return a + "a" # alt→alta
|
||||
return a + "a"
|
||||
|
||||
|
||||
def inflect_adj(lemma, gender, number):
|
||||
lemma = lemma.strip().lower()
|
||||
g = "f" if gender == "f" else "m"
|
||||
num = "SG" if number == "singular" else "PL"
|
||||
d = _ADJS.get(lemma)
|
||||
if d:
|
||||
form = d.get((g, num))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
sg = d.get((g, "SG")) or d.get(("m", "SG")) or lemma
|
||||
if num == "PL":
|
||||
pl, ok = _rule_plural(sg, g)
|
||||
return pl, ("rule" if ok else "fallback")
|
||||
return sg, "lexicon"
|
||||
# rule fallback
|
||||
base = lemma if g == "m" else _fem_of(lemma)
|
||||
if num == "SG":
|
||||
return base, "rule"
|
||||
pl, ok = _rule_plural(base, g)
|
||||
return pl, ("rule" if ok else "fallback")
|
||||
|
||||
|
||||
def lexicon_stats():
|
||||
return {
|
||||
"verb_source": "UniMorph Catalan (github.com/unimorph/cat) + kaikki.org "
|
||||
"irregulars (anar/fer/auxiliaries)",
|
||||
"noun_adj_source": "kaikki.org Catalan (Wiktionary extract)",
|
||||
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
|
||||
"unimorph_verb_forms": len(_VERBS),
|
||||
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
|
||||
"irregular_verb_lemmas": len([k for k in _IRREGV if not k.startswith("_")]),
|
||||
"participle_lemmas": len(_PART),
|
||||
"gerund_lemmas": len(_GER),
|
||||
"noun_lemmas": len(_NOUNS),
|
||||
"adj_lemmas": len(_ADJS),
|
||||
}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
|
||||
tests = [
|
||||
("cantar", "ind", "present", "first", "singular", "canto"),
|
||||
("cantar", "ind", "present", "third", "plural", "canten"),
|
||||
("ser", "ind", "present", "third", "singular", "és"),
|
||||
("haver", "ind", "present", "first", "singular", "he"),
|
||||
("anar", "ind", "present", "first", "singular", "vaig"),
|
||||
("fer", "ind", "present", "third", "singular", "fa"),
|
||||
("perdre", "ind", "present", "first", "singular", "perdo"),
|
||||
("dormir", "ind", "present", "third", "plural", "dormen"),
|
||||
("cantar", "ind", "future", "first", "singular", "cantaré"),
|
||||
("cantar", "ind", "preterite", "third", "singular", "cantà"),
|
||||
("tenir", "sbjv", "present", "first", "singular", "tingui"),
|
||||
]
|
||||
ok = 0
|
||||
for lemma, mood, tense, per, num, exp in tests:
|
||||
got, conf = conjugate(lemma, mood, tense, per, num)
|
||||
flag = "OK " if got == exp else "XX "
|
||||
ok += got == exp
|
||||
print(f" {flag}{lemma:8} {mood}/{tense:11} {per[:3]}.{num[:2]} -> {got:10} ({conf}) exp={exp}")
|
||||
print(f"verb tests {ok}/{len(tests)}")
|
||||
print(" peri-pret anar: 1sg=", peri_pret_aux("first", "singular"),
|
||||
"3pl=", peri_pret_aux("third", "plural"))
|
||||
print(" gender casa=", noun_gender("casa"), "home=", noun_gender("home"),
|
||||
"cavall=", noun_gender("cavall"), "cançó=", noun_gender("cançó"))
|
||||
print(" plural casa->", inflect_noun("casa", "plural"),
|
||||
"| plaça->", inflect_noun("plaça", "plural"),
|
||||
"| peix->", inflect_noun("peix", "plural"),
|
||||
"| braç->", inflect_noun("braç", "plural"),
|
||||
"| home->", inflect_noun("home", "plural"))
|
||||
print(" adj: alt/f/sg->", inflect_adj("alt", "f", "singular"),
|
||||
"| bonic/f/pl->", inflect_adj("bonic", "f", "plural"),
|
||||
"| vermell/f/sg->", inflect_adj("vermell", "f", "singular"))
|
||||
print(" part: cantar/f/sg->", participle("cantar", "f", "singular"),
|
||||
"| veure/f/pl->", participle("veure", "f", "plural"),
|
||||
"| fer/m/sg->", participle("fer", "m", "singular"))
|
||||
print(" ger: fer->", gerund("fer"), "| cantar->", gerund("cantar"))
|
||||
@@ -0,0 +1,423 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""morphology_de_full.py — production German morphological generator.
|
||||
|
||||
Real data, no toy tables:
|
||||
|
||||
PRIMARY — UniMorph German (github.com/unimorph/deu, CC-BY-SA 3.0).
|
||||
~219k noun forms, ~199k verb forms. Supplies:
|
||||
nouns : gender (MASC/FEM/NEUT) + case×number paradigm
|
||||
(N;NOM/ACC/DAT/GEN; MASC/FEM/NEUT; SG/PL) — the genitive -(e)s,
|
||||
dative-plural -n and the five plural classes are REAL forms, not
|
||||
guessed.
|
||||
verbs : full finite paradigm IND;{SG,PL};{1,2,3};{PRS,PST}, the past
|
||||
participle (V.PTCP;PST, incl. reattached separable prefix
|
||||
'zugefügt'), and — crucially for V2 — the SEPARATED finite form
|
||||
UniMorph records directly ('füge zu', 'steht auf').
|
||||
adjs : comparative / superlative (ADJ;CMPR, ADJ;SPRL).
|
||||
|
||||
SECONDARY — kaikki.org German (Wiktionary, CC-BY-SA/GFDL). Gap-fills noun
|
||||
gender + plural where UniMorph is thin. Never overrides UniMorph.
|
||||
|
||||
Rule fallbacks (flagged 'rule'/'fallback') for lemmas absent from both lexicons:
|
||||
present : -e/-st/-t/-en/-t/-en with e-epenthesis after -t/-d/-chn stems
|
||||
plural : gender heuristic (fem -> -(e)n, else -e / umlaut left to lexicon)
|
||||
ppart : weak ge-…-t
|
||||
Adjective ENDINGS are rule-computed by the realizer (regular closed table);
|
||||
this module only supplies the comparative/superlative STEM.
|
||||
|
||||
Perfect auxiliary (haben vs sein): sein for a curated set of intransitive
|
||||
motion / change-of-state verbs (real German lexical property), else haben.
|
||||
|
||||
Public API:
|
||||
noun_gender(lemma) -> 'm'|'f'|'n'
|
||||
decline_noun(lemma, case, number) -> (form, conf)
|
||||
pluralize(lemma) -> (form, conf)
|
||||
finite(lemma, tense, person, number) -> (form, conf) # may contain ' prefix'
|
||||
nonfinite(lemma, req) -> (form, conf) # req: 'inf'|'ppart'
|
||||
past_participle(lemma) -> (form, conf)
|
||||
separable_prefix(lemma) -> str|None
|
||||
perfect_aux(lemma) -> 'haben'|'sein'
|
||||
comparative(lemma)/superlative(lemma) -> (stem, conf)
|
||||
lexicon_stats() -> dict
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import pickle
|
||||
|
||||
_HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
_UNIMORPH = os.path.join(_HERE, "data", "deu.unimorph")
|
||||
_KAIKKI = os.path.join(_HERE, "data", "kaikki_de.jsonl")
|
||||
_CACHE = os.path.join(_HERE, "data", "de_morph_cache.pkl")
|
||||
|
||||
_GENDER = {"MASC": "m", "FEM": "f", "NEUT": "n"}
|
||||
|
||||
# intransitive motion / change-of-state verbs that take SEIN in the perfect
|
||||
_SEIN = {"gehen", "kommen", "fahren", "laufen", "rennen", "reisen", "fallen",
|
||||
"steigen", "sinken", "wachsen", "sterben", "geschehen", "passieren",
|
||||
"werden", "bleiben", "sein", "aufstehen", "einschlafen", "aufwachen",
|
||||
"ankommen", "abfahren", "aufsteigen", "erscheinen", "verschwinden",
|
||||
"fliegen", "schwimmen", "springen", "begegnen", "folgen", "gelingen",
|
||||
"wandern", "ziehen", "flüchten", "eintreten", "einsteigen", "aussteigen"}
|
||||
|
||||
|
||||
# hardcoded high-frequency irregular / auxiliary / modal paradigms (closed class,
|
||||
# verified) — consulted before the lexicon so aux+modal chains are always correct.
|
||||
_CORE = {
|
||||
"sein": {"prs": {("first", "singular"): "bin", ("second", "singular"): "bist",
|
||||
("third", "singular"): "ist", ("first", "plural"): "sind",
|
||||
("second", "plural"): "seid", ("third", "plural"): "sind"},
|
||||
"pst": {("first", "singular"): "war", ("second", "singular"): "warst",
|
||||
("third", "singular"): "war", ("first", "plural"): "waren",
|
||||
("second", "plural"): "wart", ("third", "plural"): "waren"},
|
||||
"ppart": "gewesen"},
|
||||
"haben": {"prs": {("first", "singular"): "habe", ("second", "singular"): "hast",
|
||||
("third", "singular"): "hat", ("first", "plural"): "haben",
|
||||
("second", "plural"): "habt", ("third", "plural"): "haben"},
|
||||
"pst": {("first", "singular"): "hatte", ("second", "singular"): "hattest",
|
||||
("third", "singular"): "hatte", ("first", "plural"): "hatten",
|
||||
("second", "plural"): "hattet", ("third", "plural"): "hatten"},
|
||||
"ppart": "gehabt"},
|
||||
"werden": {"prs": {("first", "singular"): "werde", ("second", "singular"): "wirst",
|
||||
("third", "singular"): "wird", ("first", "plural"): "werden",
|
||||
("second", "plural"): "werdet", ("third", "plural"): "werden"},
|
||||
"pst": {("first", "singular"): "wurde", ("second", "singular"): "wurdest",
|
||||
("third", "singular"): "wurde", ("first", "plural"): "wurden",
|
||||
("second", "plural"): "wurdet", ("third", "plural"): "wurden"},
|
||||
"ppart": "geworden"},
|
||||
}
|
||||
_MODAL_PRS = {
|
||||
"können": ("kann", "kannst", "kann", "können", "könnt", "können"),
|
||||
"müssen": ("muss", "musst", "muss", "müssen", "müsst", "müssen"),
|
||||
"wollen": ("will", "willst", "will", "wollen", "wollt", "wollen"),
|
||||
"sollen": ("soll", "sollst", "soll", "sollen", "sollt", "sollen"),
|
||||
"dürfen": ("darf", "darfst", "darf", "dürfen", "dürft", "dürfen"),
|
||||
"mögen": ("mag", "magst", "mag", "mögen", "mögt", "mögen"),
|
||||
}
|
||||
_MODAL_PST = {
|
||||
"können": ("konnte", "konntest", "konnte", "konnten", "konntet", "konnten"),
|
||||
"müssen": ("musste", "musstest", "musste", "mussten", "musstet", "mussten"),
|
||||
"wollen": ("wollte", "wolltest", "wollte", "wollten", "wolltet", "wollten"),
|
||||
"sollen": ("sollte", "solltest", "sollte", "sollten", "solltet", "sollten"),
|
||||
"dürfen": ("durfte", "durftest", "durfte", "durften", "durftet", "durften"),
|
||||
"mögen": ("mochte", "mochtest", "mochte", "mochten", "mochtet", "mochten"),
|
||||
}
|
||||
_PN_ORDER = [("first", "singular"), ("second", "singular"), ("third", "singular"),
|
||||
("first", "plural"), ("second", "plural"), ("third", "plural")]
|
||||
_MODAL_PPART = {"können": "gekonnt", "müssen": "gemusst", "wollen": "gewollt",
|
||||
"sollen": "gesollt", "dürfen": "gedurft", "mögen": "gemocht"}
|
||||
for _m, _forms in _MODAL_PRS.items():
|
||||
_CORE[_m] = {"prs": dict(zip(_PN_ORDER, _forms)),
|
||||
"pst": dict(zip(_PN_ORDER, _MODAL_PST[_m])),
|
||||
"ppart": _MODAL_PPART[_m]}
|
||||
|
||||
|
||||
def _person_num(tags):
|
||||
p = n = None
|
||||
for t in tags:
|
||||
if t in ("1", "2", "3"):
|
||||
p = {"1": "first", "2": "second", "3": "third"}[t]
|
||||
elif t == "SG":
|
||||
n = "singular"
|
||||
elif t == "PL":
|
||||
n = "plural"
|
||||
return p, n
|
||||
|
||||
|
||||
def _build_from_unimorph():
|
||||
nouns, verbs, adjs = {}, {}, {}
|
||||
if not os.path.exists(_UNIMORPH):
|
||||
return nouns, verbs, adjs
|
||||
with open(_UNIMORPH, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
line = line.rstrip("\n")
|
||||
if not line or "\t" not in line:
|
||||
continue
|
||||
parts = line.split("\t")
|
||||
if len(parts) != 3:
|
||||
continue
|
||||
lemma, form, tagstr = parts
|
||||
tags = tagstr.split(";")
|
||||
head = tags[0]
|
||||
tset = set(tags)
|
||||
if head == "N":
|
||||
rec = nouns.setdefault(lemma, {"g": None, "cases": {}, "pl": None})
|
||||
g = next((_GENDER[t] for t in tags if t in _GENDER), None)
|
||||
if g and not rec["g"]:
|
||||
rec["g"] = g
|
||||
case = next((t for t in tags if t in ("NOM", "ACC", "DAT", "GEN")), None)
|
||||
num = "plural" if "PL" in tset else ("singular" if "SG" in tset else None)
|
||||
if case and num:
|
||||
rec["cases"].setdefault((case, num), form)
|
||||
if case == "NOM" and num == "plural" and not rec["pl"]:
|
||||
rec["pl"] = form
|
||||
elif head.startswith("V"):
|
||||
rec = verbs.setdefault(lemma, {"prs": {}, "pst": {}, "ppart": None})
|
||||
if "PTCP" in head and "PST" in tset:
|
||||
rec["ppart"] = rec["ppart"] or form
|
||||
elif "IND" in tset and ("PRS" in tset or "PST" in tset):
|
||||
p, n = _person_num(tags)
|
||||
if p and n:
|
||||
slot = "prs" if "PRS" in tset else "pst"
|
||||
rec[slot].setdefault((p, n), form)
|
||||
elif head == "ADJ":
|
||||
rec = adjs.setdefault(lemma, {})
|
||||
if "CMPR" in tset:
|
||||
rec.setdefault("cmpr", form.replace("am ", "").strip())
|
||||
elif "SPRL" in tset:
|
||||
rec.setdefault("sprl", form.replace("am ", "").replace("sten", "st")
|
||||
if form.endswith("sten") else form.replace("am ", ""))
|
||||
return nouns, verbs, adjs
|
||||
|
||||
|
||||
def _build_from_kaikki(nouns):
|
||||
"""Gap-fill noun gender + plural from kaikki German."""
|
||||
if not os.path.exists(_KAIKKI):
|
||||
return
|
||||
_g = {"masculine": "m", "feminine": "f", "neuter": "n", "m": "m", "f": "f", "n": "n"}
|
||||
with open(_KAIKKI, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
try:
|
||||
d = json.loads(line)
|
||||
except Exception:
|
||||
continue
|
||||
if d.get("pos") != "noun":
|
||||
continue
|
||||
w = d.get("word", "")
|
||||
if not w or not w[0].isalpha() or " " in w:
|
||||
continue
|
||||
rec = nouns.setdefault(w, {"g": None, "cases": {}, "pl": None})
|
||||
# GENDER: Wiktionary gender is hand-curated and OVERRIDES UniMorph's
|
||||
# auto-tagged gender, which has known errors (e.g. UniMorph deu mis-
|
||||
# records Zeit=MASC, Wagen=NEUT; Wiktionary has f, m correctly).
|
||||
for h in d.get("head_templates", []) or []:
|
||||
a = h.get("args", {}) or {}
|
||||
raw = a.get("1") or a.get("g") or ""
|
||||
code = str(raw).split(",")[0].strip().lower()
|
||||
if code in _g:
|
||||
rec["g"] = _g[code]
|
||||
break
|
||||
if not rec["pl"]:
|
||||
for f in d.get("forms", []) or []:
|
||||
t = set(f.get("tags", []) or [])
|
||||
if "plural" in t and f.get("form") and "genitive" not in t:
|
||||
rec["pl"] = f["form"]
|
||||
break
|
||||
|
||||
|
||||
def _build_cache():
|
||||
nouns, verbs, adjs = _build_from_unimorph()
|
||||
_build_from_kaikki(nouns)
|
||||
data = {"nouns": nouns, "verbs": verbs, "adjs": adjs}
|
||||
try:
|
||||
with open(_CACHE, "wb") as fh:
|
||||
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
|
||||
except OSError:
|
||||
pass
|
||||
return data
|
||||
|
||||
|
||||
def _load():
|
||||
if os.path.exists(_CACHE):
|
||||
srcs = [p for p in (_UNIMORPH, _KAIKKI) if os.path.exists(p)]
|
||||
newest = max((os.path.getmtime(p) for p in srcs), default=0)
|
||||
if os.path.getmtime(_CACHE) >= newest:
|
||||
try:
|
||||
with open(_CACHE, "rb") as fh:
|
||||
return pickle.load(fh)
|
||||
except Exception:
|
||||
pass
|
||||
return _build_cache()
|
||||
|
||||
|
||||
_LEX = _load()
|
||||
_NOUNS, _VERBS, _ADJS = _LEX["nouns"], _LEX["verbs"], _LEX["adjs"]
|
||||
|
||||
|
||||
# ── nouns ────────────────────────────────────────────────────────────────────────
|
||||
def noun_gender(lemma):
|
||||
rec = _NOUNS.get(lemma) or _NOUNS.get(lemma.capitalize())
|
||||
if rec and rec.get("g"):
|
||||
return rec["g"]
|
||||
# last-resort rule: -ung/-heit/-keit/-schaft/-tät/-ion -> f ; -chen/-lein -> n
|
||||
low = lemma.lower()
|
||||
if low.endswith(("ung", "heit", "keit", "schaft", "tät", "ion", "ik", "ei")):
|
||||
return "f"
|
||||
if low.endswith(("chen", "lein", "ment", "um")):
|
||||
return "n"
|
||||
return "m"
|
||||
|
||||
|
||||
def pluralize(lemma):
|
||||
rec = _NOUNS.get(lemma) or _NOUNS.get(lemma.capitalize())
|
||||
if rec and rec.get("pl"):
|
||||
return rec["pl"], "lexicon"
|
||||
g = noun_gender(lemma)
|
||||
if g == "f":
|
||||
return (lemma + "en" if not lemma.endswith("e") else lemma + "n"), "rule"
|
||||
return (lemma if lemma.endswith(("er", "en", "el")) else lemma + "e"), "rule"
|
||||
|
||||
|
||||
def decline_noun(lemma, case, number):
|
||||
"""case in NOM/ACC/DAT/GEN, number in singular/plural."""
|
||||
rec = _NOUNS.get(lemma) or _NOUNS.get(lemma.capitalize())
|
||||
if case == "DAT" and number == "singular":
|
||||
# modern German drops the archaic dative -e ('dem Kinde' -> 'dem Kind');
|
||||
# the article carries the case. Keep bare nominative form.
|
||||
base = (rec or {}).get("cases", {}).get(("NOM", "singular")) or lemma
|
||||
return base, ("lexicon" if rec else "rule")
|
||||
if rec and rec.get("cases", {}).get((case, number)):
|
||||
return rec["cases"][(case, number)], "lexicon"
|
||||
if number == "plural":
|
||||
pl, c = pluralize(lemma)
|
||||
if case == "DAT" and not pl.endswith("n") and not pl.endswith("s"):
|
||||
return pl + "n", c # dative plural -n
|
||||
return pl, c
|
||||
# singular
|
||||
g = noun_gender(lemma)
|
||||
if case == "GEN" and g in ("m", "n"):
|
||||
return (lemma + "es" if lemma.endswith(("s", "ß", "z", "x")) else lemma + "s"), "rule"
|
||||
return lemma, "lexicon" if rec else "rule"
|
||||
|
||||
|
||||
# ── verbs ──────────────────────────────────────────────────────────────────────--
|
||||
_PRS_ENDINGS = {("first", "singular"): "e", ("second", "singular"): "st",
|
||||
("third", "singular"): "t", ("first", "plural"): "en",
|
||||
("second", "plural"): "t", ("third", "plural"): "en"}
|
||||
|
||||
|
||||
def _stem(lemma):
|
||||
if lemma.endswith("en"):
|
||||
return lemma[:-2]
|
||||
if lemma.endswith("n"):
|
||||
return lemma[:-1]
|
||||
return lemma
|
||||
|
||||
|
||||
def separable_prefix(lemma):
|
||||
"""Return the separable prefix if the lemma is a separable-prefix verb."""
|
||||
rec = _VERBS.get(lemma)
|
||||
if rec:
|
||||
for (_p, _n), form in rec.get("prs", {}).items():
|
||||
if " " in form:
|
||||
return form.rsplit(" ", 1)[1]
|
||||
_SEP = ("auf", "aus", "ab", "an", "ein", "mit", "nach", "vor", "zu", "zurück",
|
||||
"weg", "hin", "her", "los", "bei", "fest", "fort", "um", "zusammen")
|
||||
_INSEP = ("be", "ge", "er", "ver", "zer", "ent", "emp", "miss")
|
||||
for p in sorted(_SEP, key=len, reverse=True):
|
||||
if lemma.startswith(p) and len(lemma) > len(p) + 2 \
|
||||
and not lemma.startswith(_INSEP):
|
||||
return p
|
||||
return None
|
||||
|
||||
|
||||
def finite(lemma, tense, person, number):
|
||||
"""Present/past finite. For separable verbs the returned string is the
|
||||
UniMorph SEPARATED form 'stem prefix' (realizer places prefix per V2)."""
|
||||
slot = "prs" if tense == "present" else "pst"
|
||||
if lemma in _CORE and _CORE[lemma].get(slot, {}).get((person, number)):
|
||||
return _CORE[lemma][slot][(person, number)], "lexicon"
|
||||
rec = _VERBS.get(lemma)
|
||||
if rec and rec.get(slot, {}).get((person, number)):
|
||||
return rec[slot][(person, number)], "lexicon"
|
||||
# rule fallback (present only reliable; past weak -te)
|
||||
stem = _stem(lemma)
|
||||
pref = separable_prefix(lemma)
|
||||
if pref:
|
||||
stem = _stem(lemma[len(pref):])
|
||||
if tense == "present":
|
||||
end = _PRS_ENDINGS[(person, number)]
|
||||
if stem.endswith(("t", "d", "chn", "ffn", "gn")) and end in ("st", "t"):
|
||||
end = "e" + end
|
||||
form = stem + end
|
||||
else:
|
||||
form = stem + ("ete" if stem.endswith(("t", "d")) else "te")
|
||||
if (person, number) == ("second", "singular"):
|
||||
form += "st"
|
||||
elif number == "plural" and person != "second":
|
||||
form += "n"
|
||||
elif (person, number) == ("second", "plural"):
|
||||
form += "t"
|
||||
if pref:
|
||||
return f"{form} {pref}", "rule"
|
||||
return form, "rule"
|
||||
|
||||
|
||||
def _weak_t(stem):
|
||||
return stem + ("et" if stem.endswith(("t", "d", "chn", "ffn", "gn")) else "t")
|
||||
|
||||
|
||||
def past_participle(lemma):
|
||||
if lemma in _CORE:
|
||||
return _CORE[lemma]["ppart"], "lexicon"
|
||||
rec = _VERBS.get(lemma)
|
||||
if rec and rec.get("ppart"):
|
||||
return rec["ppart"], "lexicon"
|
||||
stem = _stem(lemma)
|
||||
pref = separable_prefix(lemma)
|
||||
_INSEP = ("be", "ge", "er", "ver", "zer", "ent", "emp", "miss")
|
||||
if pref:
|
||||
inner = _stem(lemma[len(pref):])
|
||||
return pref + "ge" + _weak_t(inner), "rule"
|
||||
if lemma.startswith(_INSEP):
|
||||
return _weak_t(stem), "rule"
|
||||
return "ge" + _weak_t(stem), "rule"
|
||||
|
||||
|
||||
def nonfinite(lemma, req):
|
||||
if req == "ppart":
|
||||
return past_participle(lemma)
|
||||
return lemma, "lexicon" if lemma in _VERBS else "rule" # infinitive
|
||||
|
||||
|
||||
def perfect_aux(lemma):
|
||||
return "sein" if lemma in _SEIN else "haben"
|
||||
|
||||
|
||||
# ── adjectives ────────────────────────────────────────────────────────────────---
|
||||
_ADJ_IRREG_SPRL = {"gut": "best", "groß": "größt", "hoch": "höchst",
|
||||
"nah": "nächst", "viel": "meist", "gern": "liebst"}
|
||||
|
||||
|
||||
def comparative(lemma):
|
||||
rec = _ADJS.get(lemma)
|
||||
if rec and rec.get("cmpr"):
|
||||
return rec["cmpr"], "lexicon"
|
||||
return lemma + "er", "rule"
|
||||
|
||||
|
||||
def superlative(lemma):
|
||||
"""Return the bare superlative STEM (realizer adds 'am ...en' or '-e' ending)."""
|
||||
if lemma in _ADJ_IRREG_SPRL:
|
||||
return _ADJ_IRREG_SPRL[lemma], "lexicon"
|
||||
# derive from the comparative so umlaut is carried (alt->älter->ältest)
|
||||
cmpr, cconf = comparative(lemma)
|
||||
base = cmpr[:-2] if cmpr.endswith("er") else lemma
|
||||
end = "est" if base.endswith(("t", "d", "s", "ß", "z", "sch")) else "st"
|
||||
return base + end, cconf
|
||||
|
||||
|
||||
def lexicon_stats():
|
||||
return {
|
||||
"source": "UniMorph deu (primary) + kaikki.org German (gap-fill gender/plural)",
|
||||
"license": "CC-BY-SA 3.0 (UniMorph); CC-BY-SA/GFDL (Wiktionary)",
|
||||
"noun_lemmas": len(_NOUNS),
|
||||
"nouns_with_gender": sum(1 for v in _NOUNS.values() if v.get("g")),
|
||||
"nouns_with_plural": sum(1 for v in _NOUNS.values() if v.get("pl")),
|
||||
"verb_lemmas": len(_VERBS),
|
||||
"verbs_with_ppart": sum(1 for v in _VERBS.values() if v.get("ppart")),
|
||||
"adj_lemmas": len(_ADJS),
|
||||
}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
|
||||
for w in ("Hund", "Frau", "Kind", "Mann", "Buch", "Blume"):
|
||||
print(f" {w}: gender={noun_gender(w)} pl={pluralize(w)} "
|
||||
f"gen.sg={decline_noun(w, 'GEN', 'singular')} "
|
||||
f"dat.pl={decline_noun(w, 'DAT', 'plural')}")
|
||||
for v in ("machen", "gehen", "aufstehen", "sein", "haben", "arbeiten"):
|
||||
print(f" {v}: 3sg.prs={finite(v, 'present', 'third', 'singular')} "
|
||||
f"3sg.pst={finite(v, 'past', 'third', 'singular')} "
|
||||
f"ppart={past_participle(v)} aux={perfect_aux(v)} sep={separable_prefix(v)}")
|
||||
for a in ("schnell", "gut", "groß", "alt"):
|
||||
print(f" {a}: cmpr={comparative(a)} sprl={superlative(a)}")
|
||||
@@ -0,0 +1,562 @@
|
||||
"""morphology_es_full.py — production-grade Spanish morphological generator.
|
||||
|
||||
NOT a toy. Backed by a real, broad, licensed lexicon:
|
||||
|
||||
UniMorph Spanish (github.com/unimorph/spa, CC-BY-SA 3.0, Wiktionary-derived)
|
||||
1,196,245 inflected forms:
|
||||
6,695 verb lemmas — full paradigms: indicative (present/preterite/
|
||||
imperfect/future), conditional, present & imperfect
|
||||
subjunctive, affirmative imperative, formal/informal
|
||||
48,353 noun lemmas — WITH inherent gender (N;FEM/MASC;SG/PL)
|
||||
16,984 adj lemmas — gender + number paradigms
|
||||
|
||||
Fallbacks (so we degrade, never crash, on out-of-vocabulary input):
|
||||
- verbs : mlconjug3 (ML paradigm model, conjugates ANY Spanish verb) then a
|
||||
hand-rolled regular-ending generator
|
||||
- nouns : gender heuristic (endings) + regular pluralization
|
||||
- adjs : -o/-a gender rule + regular pluralization
|
||||
|
||||
Every generated form carries a CONFIDENCE flag:
|
||||
"lexicon" form came straight from UniMorph (trust: high)
|
||||
"model" form came from mlconjug3 (trust: high)
|
||||
"rule" form came from a deterministic rule (trust: medium)
|
||||
"fallback" we could not inflect; returned lemma as-is (trust: low → FLAG)
|
||||
|
||||
Public API (used by realizer_es.py):
|
||||
conjugate(lemma, mood, tense, person, number, formality="informal") -> (form, conf)
|
||||
participle(lemma) -> (form, conf) # past participle (compound tenses)
|
||||
gerund(lemma) -> (form, conf)
|
||||
noun_gender(lemma) -> "m"|"f"
|
||||
inflect_noun(lemma, number) -> (form, conf)
|
||||
inflect_adj(lemma, gender, number) -> (form, conf)
|
||||
attach_enclitics(verb_form, clitics) -> str # accent-correct enclisis
|
||||
lexicon_stats() -> dict
|
||||
"""
|
||||
import os
|
||||
import pickle
|
||||
import unicodedata
|
||||
|
||||
_HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
_UNIMORPH = os.path.join(_HERE, "data", "spa.unimorph")
|
||||
_CACHE = os.path.join(_HERE, "data", "es_morph_cache.pkl")
|
||||
|
||||
# ── canonical feature keys the realizer speaks, mapped to UniMorph tags ─────────
|
||||
# mood/tense pair -> the UniMorph feature substring that identifies it
|
||||
_VERB_KEYMAP = {
|
||||
("ind", "present"): ("IND", "PRS", None),
|
||||
("ind", "preterite"): ("IND", "PST", "PFV"),
|
||||
("ind", "imperfect"): ("IND", "PST", "IPFV"),
|
||||
("ind", "future"): ("IND", "FUT", None),
|
||||
("ind", "conditional"):("COND", None, None),
|
||||
("sbjv", "present"): ("SBJV", "PRS", None),
|
||||
("sbjv", "imperfect"): ("SBJV", "PST", "LGSPEC1"), # -ra form
|
||||
("imp", "present"): ("POS", "IMP", None),
|
||||
}
|
||||
_PERSON = {"first": "1", "second": "2", "third": "3"}
|
||||
_NUMBER = {"singular": "SG", "plural": "PL"}
|
||||
|
||||
|
||||
# ── build / load the compact lexicon ───────────────────────────────────────────
|
||||
def _feat_set(tag):
|
||||
return set(tag.split(";"))
|
||||
|
||||
|
||||
def _build_cache():
|
||||
verbs = {} # (lemma, canonkey) -> form canonkey e.g. "ind|present|1|SG|infm"
|
||||
nouns = {} # lemma -> {"g": "m"/"f", "SG": form, "PL": form}
|
||||
adjs = {} # lemma -> {("m","SG"): form, ...}
|
||||
part = {} # lemma -> masc-sg participle
|
||||
ger = {} # lemma -> gerund
|
||||
|
||||
with open(_UNIMORPH, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
line = line.rstrip("\n")
|
||||
if not line or "\t" not in line:
|
||||
continue
|
||||
parts = line.split("\t")
|
||||
if len(parts) != 3:
|
||||
continue
|
||||
lemma, form, tag = parts
|
||||
f = _feat_set(tag)
|
||||
head = tag.split(";")[0]
|
||||
|
||||
if head == "V":
|
||||
# skip clitic-bearing rows (we generate clitics ourselves)
|
||||
if "PRO" in f:
|
||||
continue
|
||||
if "V.PTCP" in f and "PST" in f and "MASC" in f and "SG" in f:
|
||||
part.setdefault(lemma, form)
|
||||
continue
|
||||
if "V.CVB" in f or "NFIN" in f or "V.PTCP" in f:
|
||||
if "V.CVB" in f:
|
||||
ger.setdefault(lemma, form)
|
||||
continue
|
||||
# identify mood/tense
|
||||
mt = None
|
||||
for (mood, tense), (a, b, c) in _VERB_KEYMAP.items():
|
||||
if a not in f:
|
||||
continue
|
||||
if b is not None and b not in f:
|
||||
continue
|
||||
if c is not None and c not in f:
|
||||
continue
|
||||
# disambiguate IND;PST needing PFV vs IPFV
|
||||
if a == "IND" and b == "PST" and c not in f:
|
||||
continue
|
||||
mt = (mood, tense)
|
||||
break
|
||||
if mt is None:
|
||||
continue
|
||||
person = next((p for p in ("1", "2", "3") if p in f), None)
|
||||
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
|
||||
if person is None or number is None:
|
||||
continue
|
||||
formal = "form" if "FORM" in f else ("infm" if "INFM" in f else "any")
|
||||
key = f"{mt[0]}|{mt[1]}|{person}|{number}|{formal}"
|
||||
verbs.setdefault((lemma, key), form)
|
||||
|
||||
elif head == "N":
|
||||
# substring test handles epicene "MASC+FEM" (-> masc citation)
|
||||
g = "m" if "MASC" in tag else ("f" if "FEM" in tag else None)
|
||||
num = "SG" if "SG" in f else ("PL" if "PL" in f else None)
|
||||
if num is None:
|
||||
continue
|
||||
# store forms keyed by (gender,number); animate nouns list BOTH
|
||||
# genders under one lemma (niño -> niño/niña). Resolve citation
|
||||
# gender in a post-pass (gender of the row whose form == lemma).
|
||||
d = nouns.setdefault(lemma, {})
|
||||
d.setdefault("_rows", []).append((g, num, form))
|
||||
|
||||
elif head == "ADJ":
|
||||
g = "m" if "MASC" in tag else ("f" if "FEM" in tag else "m")
|
||||
num = "SG" if "SG" in f else ("PL" if "PL" in f else None)
|
||||
if num is None:
|
||||
continue
|
||||
adjs.setdefault(lemma, {})[(g, num)] = form
|
||||
|
||||
# post-pass: resolve noun citation gender + default SG/PL forms
|
||||
for lemma, d in nouns.items():
|
||||
rows = d.pop("_rows", [])
|
||||
# citation gender = gender of the row whose form == lemma; else first MASC;
|
||||
# else first seen gender.
|
||||
cite_g = None
|
||||
for g, num, form in rows:
|
||||
if form == lemma and g:
|
||||
cite_g = g
|
||||
break
|
||||
if cite_g is None:
|
||||
for g, num, form in rows:
|
||||
if g == "m":
|
||||
cite_g = "m"
|
||||
break
|
||||
if cite_g is None:
|
||||
cite_g = next((g for g, _, _ in rows if g), "m")
|
||||
d["g"] = cite_g
|
||||
for g, num, form in rows:
|
||||
d[(g, num)] = form
|
||||
d["SG"] = d.get((cite_g, "SG")) or next((f for g, n, f in rows if n == "SG"), lemma)
|
||||
d["PL"] = d.get((cite_g, "PL")) or next((f for g, n, f in rows if n == "PL"), None)
|
||||
|
||||
# post-pass: UniMorph omits the identity inflection (masc-sg == lemma) for
|
||||
# adjectives, so fill it in; without this a fem-sg row wrongly satisfies a
|
||||
# masc-sg request (alto -> alta bug).
|
||||
for lemma, d in adjs.items():
|
||||
d.setdefault(("m", "SG"), lemma)
|
||||
|
||||
data = {"verbs": verbs, "nouns": nouns, "adjs": adjs, "part": part, "ger": ger}
|
||||
try:
|
||||
with open(_CACHE, "wb") as fh:
|
||||
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
|
||||
except OSError:
|
||||
pass
|
||||
return data
|
||||
|
||||
|
||||
def _load():
|
||||
if os.path.exists(_CACHE) and os.path.getmtime(_CACHE) >= os.path.getmtime(_UNIMORPH):
|
||||
try:
|
||||
with open(_CACHE, "rb") as fh:
|
||||
return pickle.load(fh)
|
||||
except Exception:
|
||||
pass
|
||||
return _build_cache()
|
||||
|
||||
|
||||
_LEX = _load()
|
||||
_VERBS, _NOUNS, _ADJS, _PART, _GER = (
|
||||
_LEX["verbs"], _LEX["nouns"], _LEX["adjs"], _LEX["part"], _LEX["ger"])
|
||||
|
||||
# ── mlconjug3 fallback (lazy) ───────────────────────────────────────────────────
|
||||
_MLC = None
|
||||
_MLC_TENSE = { # (mood,tense) -> (mlconjug mood label, tense label)
|
||||
("ind", "present"): ("Indicativo", "Indicativo presente"),
|
||||
("ind", "preterite"): ("Indicativo", "Indicativo pretérito perfecto simple"),
|
||||
("ind", "imperfect"): ("Indicativo", "Indicativo pretérito imperfecto"),
|
||||
("ind", "future"): ("Indicativo", "Indicativo futuro"),
|
||||
("ind", "conditional"): ("Condicional", "Condicional Condicional"),
|
||||
("sbjv", "present"): ("Subjuntivo", "Subjuntivo presente"),
|
||||
("sbjv", "imperfect"): ("Subjuntivo", "Subjuntivo pretérito imperfecto 1"),
|
||||
("imp", "present"): ("Imperativo", "Imperativo Afirmativo"),
|
||||
}
|
||||
_MLC_SLOT = { # (person,number) -> mlconjug slot key
|
||||
("first", "singular"): "1s", ("second", "singular"): "2s",
|
||||
("third", "singular"): "3s", ("first", "plural"): "1p",
|
||||
("second", "plural"): "2p", ("third", "plural"): "3p",
|
||||
}
|
||||
|
||||
|
||||
def _mlc_conjugate(lemma, mood, tense, person, number):
|
||||
global _MLC
|
||||
try:
|
||||
if _MLC is None:
|
||||
from mlconjug3 import Conjugator
|
||||
_MLC = Conjugator(language="es")
|
||||
v = _MLC.conjugate(lemma)
|
||||
if v is None:
|
||||
return None
|
||||
info = v.conjug_info
|
||||
m, t = _MLC_TENSE.get((mood, tense), (None, None))
|
||||
if m is None or m not in info or t not in info[m]:
|
||||
return None
|
||||
block = info[m][t]
|
||||
slot = _MLC_SLOT.get((person, number))
|
||||
if isinstance(block, dict) and slot in block and block[slot]:
|
||||
return block[slot]
|
||||
return None
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
# ── regular-ending rule fallback (last resort, deterministic) ───────────────────
|
||||
def _vclass(lemma):
|
||||
return lemma[-2:] if lemma[-2:] in ("ar", "er", "ir") else "ar"
|
||||
|
||||
|
||||
def _stem(lemma):
|
||||
return lemma[:-2]
|
||||
|
||||
|
||||
_REG = {
|
||||
("ind", "present", "ar"): ["o", "as", "a", "amos", "áis", "an"],
|
||||
("ind", "present", "er"): ["o", "es", "e", "emos", "éis", "en"],
|
||||
("ind", "present", "ir"): ["o", "es", "e", "imos", "ís", "en"],
|
||||
("ind", "preterite", "ar"): ["é", "aste", "ó", "amos", "asteis", "aron"],
|
||||
("ind", "preterite", "er"): ["í", "iste", "ió", "imos", "isteis", "ieron"],
|
||||
("ind", "preterite", "ir"): ["í", "iste", "ió", "imos", "isteis", "ieron"],
|
||||
("ind", "imperfect", "ar"): ["aba", "abas", "aba", "ábamos", "abais", "aban"],
|
||||
("ind", "imperfect", "er"): ["ía", "ías", "ía", "íamos", "íais", "ían"],
|
||||
("ind", "imperfect", "ir"): ["ía", "ías", "ía", "íamos", "íais", "ían"],
|
||||
("sbjv", "present", "ar"): ["e", "es", "e", "emos", "éis", "en"],
|
||||
("sbjv", "present", "er"): ["a", "as", "a", "amos", "áis", "an"],
|
||||
("sbjv", "present", "ir"): ["a", "as", "a", "amos", "áis", "an"],
|
||||
("sbjv", "imperfect", "ar"): ["ara", "aras", "ara", "áramos", "arais", "aran"],
|
||||
("sbjv", "imperfect", "er"): ["iera", "ieras", "iera", "iéramos", "ierais", "ieran"],
|
||||
("sbjv", "imperfect", "ir"): ["iera", "ieras", "iera", "iéramos", "ierais", "ieran"],
|
||||
}
|
||||
_FUT = ["é", "ás", "á", "emos", "éis", "án"]
|
||||
_COND = ["ía", "ías", "ía", "íamos", "íais", "ían"]
|
||||
|
||||
|
||||
def _slot_idx(person, number):
|
||||
base = {"first": 0, "second": 1, "third": 2}[person]
|
||||
return base + (0 if number == "singular" else 3)
|
||||
|
||||
|
||||
def _rule_conjugate(lemma, mood, tense, person, number):
|
||||
if len(lemma) < 3 or lemma[-2:] not in ("ar", "er", "ir"):
|
||||
return None
|
||||
vc, st, i = _vclass(lemma), _stem(lemma), _slot_idx(person, number)
|
||||
if tense == "future":
|
||||
return lemma + _FUT[i]
|
||||
if tense == "conditional":
|
||||
return lemma + _COND[i]
|
||||
table = _REG.get((mood, tense, vc))
|
||||
if table:
|
||||
return st + table[i]
|
||||
if mood == "imp" and tense == "present":
|
||||
# affirmative tú imperative = 3sg present indicative
|
||||
pres = _REG.get(("ind", "present", vc))
|
||||
return st + pres[2] if number == "singular" else st + pres[5]
|
||||
return None
|
||||
|
||||
|
||||
# ── PUBLIC: verb conjugation ────────────────────────────────────────────────────
|
||||
def conjugate(lemma, mood, tense, person, number, formality="informal"):
|
||||
"""Return (surface, confidence). mood in ind|sbjv|imp; tense per _VERB_KEYMAP."""
|
||||
lemma = lemma.strip().lower()
|
||||
p, n = _PERSON.get(person), _NUMBER.get(number)
|
||||
formal = "form" if formality == "formal" else "infm"
|
||||
if p and n:
|
||||
for fkey in (formal, "any", "infm" if formal == "form" else "form"):
|
||||
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}|{fkey}"))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
m = _mlc_conjugate(lemma, mood, tense, person, number)
|
||||
if m:
|
||||
return m, "model"
|
||||
r = _rule_conjugate(lemma, mood, tense, person, number)
|
||||
if r:
|
||||
return r, "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
_IRREG_PART = { # guarantee the common irregular participles
|
||||
"escribir": "escrito", "describir": "descrito", "abrir": "abierto",
|
||||
"cubrir": "cubierto", "descubrir": "descubierto", "morir": "muerto",
|
||||
"poner": "puesto", "ver": "visto", "volver": "vuelto", "devolver": "devuelto",
|
||||
"hacer": "hecho", "deshacer": "deshecho", "decir": "dicho", "romper": "roto",
|
||||
"resolver": "resuelto", "freír": "frito", "imprimir": "impreso",
|
||||
"satisfacer": "satisfecho", "prever": "previsto", "revolver": "revuelto",
|
||||
}
|
||||
|
||||
|
||||
def participle(lemma):
|
||||
lemma = lemma.strip().lower()
|
||||
if lemma in _IRREG_PART:
|
||||
return _IRREG_PART[lemma], "lexicon"
|
||||
if lemma in _PART:
|
||||
return _PART[lemma], "lexicon"
|
||||
if lemma.endswith("ar"):
|
||||
return lemma[:-2] + "ado", "rule"
|
||||
if lemma[-2:] in ("er", "ir"):
|
||||
return lemma[:-2] + "ido", "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
_IRREG_GER = {"dormir": "durmiendo", "morir": "muriendo", "pedir": "pidiendo",
|
||||
"sentir": "sintiendo", "mentir": "mintiendo", "servir": "sirviendo",
|
||||
"venir": "viniendo", "decir": "diciendo", "poder": "pudiendo",
|
||||
"ir": "yendo", "leer": "leyendo", "creer": "creyendo",
|
||||
"oír": "oyendo", "traer": "trayendo", "caer": "cayendo",
|
||||
"construir": "construyendo", "huir": "huyendo", "reír": "riendo"}
|
||||
|
||||
|
||||
def gerund(lemma):
|
||||
lemma = lemma.strip().lower()
|
||||
if lemma in _IRREG_GER:
|
||||
return _IRREG_GER[lemma], "lexicon"
|
||||
if lemma in _GER:
|
||||
return _GER[lemma], "lexicon"
|
||||
if lemma.endswith("ar"):
|
||||
return lemma[:-2] + "ando", "rule"
|
||||
if lemma[-2:] in ("er", "ir"):
|
||||
return lemma[:-2] + "iendo", "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
# ── PUBLIC: noun gender + number ────────────────────────────────────────────────
|
||||
_INVARIANT_PL = {"lunes", "martes", "miércoles", "jueves", "viernes",
|
||||
"crisis", "tesis", "análisis", "dosis", "virus", "paraguas"}
|
||||
|
||||
|
||||
def _gender_heuristic(noun):
|
||||
for suf, g in (("ión", "f"), ("dad", "f"), ("tad", "f"), ("umbre", "f"),
|
||||
("sis", "f"), ("ez", "f"), ("triz", "f"),
|
||||
("ema", "m"), ("ama", "m"), ("oma", "m"), ("aje", "m"),
|
||||
("or", "m"), ("án", "m"), ("ín", "m")):
|
||||
if noun.endswith(suf):
|
||||
return g
|
||||
if noun.endswith("o"):
|
||||
return "m"
|
||||
if noun.endswith("a"):
|
||||
return "f"
|
||||
return "m"
|
||||
|
||||
|
||||
def noun_gender(lemma):
|
||||
lemma = lemma.strip().lower()
|
||||
d = _NOUNS.get(lemma)
|
||||
if d and d.get("g"):
|
||||
return d["g"]
|
||||
return _gender_heuristic(lemma)
|
||||
|
||||
|
||||
def _regular_plural(noun):
|
||||
if noun in _INVARIANT_PL:
|
||||
return noun
|
||||
if not noun:
|
||||
return noun
|
||||
last = noun[-1]
|
||||
if last == "z":
|
||||
return noun[:-1] + "ces"
|
||||
if last in "aeiouáéíóú":
|
||||
# stressed final vowel í/ú -> +es (rubí->rubíes), else +s
|
||||
if last in "íú":
|
||||
return noun + "es"
|
||||
return noun + "s"
|
||||
if last == "s":
|
||||
# esdrújula / stress-final handled crudely; most polysyllables invariant
|
||||
return noun
|
||||
return noun + "es"
|
||||
|
||||
|
||||
def inflect_noun(lemma, number, gender=None):
|
||||
lemma = lemma.strip().lower()
|
||||
d = _NOUNS.get(lemma)
|
||||
num = "SG" if number == "singular" else "PL"
|
||||
if d:
|
||||
# honor a requested gender for animate nouns (gato -> gata)
|
||||
if gender and (gender, num) in d:
|
||||
return d[(gender, num)], "lexicon"
|
||||
if d.get(num):
|
||||
return d[num], "lexicon"
|
||||
if number == "singular":
|
||||
return lemma, "rule" if not d else "lexicon"
|
||||
return _regular_plural(lemma), "rule"
|
||||
|
||||
|
||||
# ── PUBLIC: adjective agreement ─────────────────────────────────────────────────
|
||||
_INV_GENDER_ADJ = {"español": "española", "trabajador": "trabajadora",
|
||||
"hablador": "habladora", "encantador": "encantadora",
|
||||
"alemán": "alemana", "francés": "francesa", "inglés": "inglesa"}
|
||||
|
||||
|
||||
def inflect_adj(lemma, gender, number):
|
||||
lemma = lemma.strip().lower()
|
||||
d = _ADJS.get(lemma)
|
||||
num = "SG" if number == "singular" else "PL"
|
||||
if d:
|
||||
form = d.get((gender, num))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
# gender-invariant adjective (grande, feliz, azul): fem == masc.
|
||||
# For a missing plural, pluralize this gender's singular form.
|
||||
sg = d.get((gender, "SG")) or d.get(("m", "SG")) or lemma
|
||||
if number == "plural":
|
||||
return _regular_plural(sg), "rule"
|
||||
return sg, "lexicon"
|
||||
# rule fallback
|
||||
a = lemma
|
||||
if gender == "f":
|
||||
if a in _INV_GENDER_ADJ:
|
||||
a = _INV_GENDER_ADJ[a]
|
||||
elif a.endswith("o"):
|
||||
a = a[:-1] + "a"
|
||||
if number == "plural":
|
||||
a = _regular_plural(a)
|
||||
return a, ("rule" if (a != lemma or gender == "m") else "rule")
|
||||
|
||||
|
||||
# ── PUBLIC: clitic enclisis (dá + me + lo -> dámelo) ────────────────────────────
|
||||
def _strip_accents(s):
|
||||
return "".join(c for c in unicodedata.normalize("NFD", s)
|
||||
if unicodedata.category(c) != "Mn")
|
||||
|
||||
|
||||
def _count_syllables_vowelgroups(word):
|
||||
# crude: count vowel groups
|
||||
w = _strip_accents(word).lower()
|
||||
groups, prev = 0, False
|
||||
for ch in w:
|
||||
isv = ch in "aeiou"
|
||||
if isv and not prev:
|
||||
groups += 1
|
||||
prev = isv
|
||||
return groups
|
||||
|
||||
|
||||
def _host_stress_from_end(word):
|
||||
"""Stressed-syllable index counted from the end (1=last) of a verb host."""
|
||||
syls = _count_syllables_vowelgroups(word)
|
||||
if any(c in "áéíóú" for c in word):
|
||||
return None # already carries its own accent
|
||||
if word[-2:] in ("ar", "er", "ir"): # infinitive: oxytone
|
||||
return 1
|
||||
if word.endswith("ndo"): # gerund: paroxytone
|
||||
return 2
|
||||
if word[-1:] in "aeiouns" and syls >= 2: # default paroxytone
|
||||
return 2
|
||||
return 1 # monosyllable / consonant-final oxytone
|
||||
|
||||
|
||||
def attach_enclitics(verb_form, clitics):
|
||||
"""Append clitic pronouns to a verb (imperative/infinitive/gerund enclisis)
|
||||
and add a written accent when the resulting word becomes esdrújula/
|
||||
sobreesdrújula (stress >= 3 syllables from the end): dá+me+lo -> dámelo,
|
||||
lleva+me -> llévame, but dar+te -> darte and da+me -> dame (no accent)."""
|
||||
if not clitics:
|
||||
return verb_form
|
||||
tail = "".join(clitics)
|
||||
if any(c in "áéíóú" for c in verb_form): # host already accented
|
||||
return verb_form + tail
|
||||
sfe = _host_stress_from_end(verb_form)
|
||||
total_sfe = sfe + len(clitics) # each clitic = 1 syllable
|
||||
if total_sfe >= 3:
|
||||
return _accentuate_nucleus(verb_form, sfe) + tail
|
||||
return verb_form + tail
|
||||
|
||||
|
||||
def _accentuate_nucleus(word, sfe):
|
||||
"""Put a written accent on the syllable `sfe` positions from the word's end."""
|
||||
vowels = "aeiou"
|
||||
nuclei = [i for i, ch in enumerate(word) if ch in vowels]
|
||||
if not nuclei or sfe > len(nuclei):
|
||||
return word
|
||||
i = nuclei[-sfe]
|
||||
acc = {"a": "á", "e": "é", "i": "í", "o": "ó", "u": "ú"}
|
||||
return word[:i] + acc[word[i]] + word[i + 1:]
|
||||
|
||||
|
||||
def _accentuate_last_stressed(word):
|
||||
# Restore the host's ORIGINAL lexical stress with a written accent.
|
||||
# Default Spanish stress: word ending in vowel/n/s -> penultimate syllable;
|
||||
# otherwise (e.g. infinitives in -r) -> last syllable.
|
||||
vowels = "aeiou"
|
||||
nuclei = [i for i, ch in enumerate(word) if ch in vowels]
|
||||
if not nuclei:
|
||||
return word
|
||||
if word[-1] in "aeiouns" and len(nuclei) >= 2:
|
||||
i = nuclei[-2] # paroxytone: penult nucleus
|
||||
else:
|
||||
i = nuclei[-1] # oxytone / monosyllable: last nucleus
|
||||
acc = {"a": "á", "e": "é", "i": "í", "o": "ó", "u": "ú"}
|
||||
return word[:i] + acc[word[i]] + word[i + 1:]
|
||||
|
||||
|
||||
def lexicon_stats():
|
||||
return {
|
||||
"source": "UniMorph Spanish (github.com/unimorph/spa)",
|
||||
"license": "CC-BY-SA 3.0 (Wiktionary-derived)",
|
||||
"total_forms": sum(len(v) for v in (_VERBS, _NOUNS, _ADJS)) if False else None,
|
||||
"verb_forms": len(_VERBS),
|
||||
"verb_lemmas": len({k[0] for k in _VERBS}),
|
||||
"noun_lemmas": len(_NOUNS),
|
||||
"adj_lemmas": len(_ADJS),
|
||||
"participles": len(_PART),
|
||||
"gerunds": len(_GER),
|
||||
}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
import json
|
||||
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
|
||||
tests = [
|
||||
("hablar", "ind", "present", "first", "singular", "hablo"),
|
||||
("comer", "ind", "present", "third", "plural", "comen"),
|
||||
("vivir", "ind", "present", "first", "plural", "vivimos"),
|
||||
("ser", "ind", "present", "third", "singular", "es"),
|
||||
("ir", "ind", "preterite", "first", "singular", "fui"),
|
||||
("tener", "ind", "future", "first", "singular", "tendré"),
|
||||
("hacer", "sbjv", "present", "first", "singular", "haga"),
|
||||
("dormir", "ind", "present", "first", "singular", "duermo"),
|
||||
("pensar", "sbjv", "present", "third", "singular", "piense"),
|
||||
("dar", "ind", "preterite", "third", "singular", "dio"),
|
||||
("poner", "ind", "conditional", "first", "singular", "pondría"),
|
||||
]
|
||||
ok = 0
|
||||
for lemma, mood, tense, per, num, exp in tests:
|
||||
got, conf = conjugate(lemma, mood, tense, per, num)
|
||||
flag = "OK " if got == exp else "XX "
|
||||
if got == exp:
|
||||
ok += 1
|
||||
print(f" {flag}{lemma:8} {mood}/{tense} {per[:3]}.{num[:2]:3} -> {got:14} ({conf}) exp={exp}")
|
||||
print(f"verb tests {ok}/{len(tests)}")
|
||||
print(" gender casa:", noun_gender("casa"), "| problema:", noun_gender("problema"),
|
||||
"| agua:", noun_gender("agua"), "| mano:", noun_gender("mano"))
|
||||
print(" plural: luz->", inflect_noun("luz", "plural"), "| rey->", inflect_noun("rey", "plural"))
|
||||
print(" adj: rojo/f/pl->", inflect_adj("rojo", "f", "plural"),
|
||||
"| feliz/m/pl->", inflect_adj("feliz", "m", "plural"),
|
||||
"| grande/f/pl->", inflect_adj("grande", "f", "plural"))
|
||||
print(" enclisis: da+[me,lo]->", attach_enclitics("da", ["me", "lo"]),
|
||||
"| di+[me]->", attach_enclitics("di", ["me"]),
|
||||
"| dar+[se,lo]->", attach_enclitics("dar", ["se", "lo"]))
|
||||
@@ -0,0 +1,629 @@
|
||||
"""morphology_fr_full.py — production-grade French morphological generator.
|
||||
|
||||
Same architecture as morphology_it_full.py (shared Romance engine); French-specific
|
||||
data and rules swapped in. Backed by three real, Wiktionary-lineage sources:
|
||||
|
||||
VERBS
|
||||
UniMorph French (github.com/unimorph/fra, CC-BY-SA 3.0)
|
||||
7,535 verb lemmas × full paradigm, CLEAN orthography:
|
||||
indicatif présent / imparfait (PST;IPFV) / passé simple (PST;PFV) /
|
||||
futur, conditionnel (COND), subjonctif présent (SBJV;PRS) /
|
||||
subjonctif imparfait (SBJV;PST), impératif (POS;IMP), infinitif (NFIN),
|
||||
participe présent (V.CVB/V.PTCP;PRS), participe passé (V.PTCP;PST, m.sg).
|
||||
fr_irreg_verbs.json — high-frequency verbs UniMorph MISSES or mis-slots,
|
||||
above all ÊTRE (absent from UniMorph fra), plus avoir/aller/faire/… — the
|
||||
auxiliaries the passé-composé + être-agreement system depends on. Extracted
|
||||
from kaikki.org French (build_fr_irreg.py), reflexive/multiword forms
|
||||
dropped. This layer takes PRIORITY.
|
||||
|
||||
NOUNS + ADJECTIVES — kaikki.org French (Wiktionary extract, CC-BY-SA 3.0)
|
||||
noun lemmas WITH inherent gender (head-template arg) + real plural
|
||||
(cheval->chevaux, œil->yeux, invariable -s/-x/-z), resolved PER LEMMA.
|
||||
adjective lemmas with real feminine + plural (petit->petite/petits/petites,
|
||||
beau->belle/beaux/belles, heureux->heureuse, rouge invariant-gender).
|
||||
|
||||
Fallbacks (degrade, never crash, on OOV input):
|
||||
verbs : rule generator for -er / -ir(-iss-) / -re (with -cer/-ger spelling,
|
||||
future/conditional stems, imparfait/subjonctif endings)
|
||||
nouns : gender heuristic (endings) + rule pluralization (-al->-aux, -eau->-eaux)
|
||||
adjs : fem/plural agreement rules (-er->-ère, -eux->-euse, -f->-ve, +e default)
|
||||
|
||||
Confidence flag on every form: "lexicon" | "rule" | "fallback".
|
||||
|
||||
Public API (used by realizer_fr.py): identical signature to morphology_it_full.
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import pickle
|
||||
|
||||
_HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
_UNIMORPH = os.path.join(_HERE, "data", "fra.unimorph")
|
||||
_IRREG = os.path.join(_HERE, "data", "fr_irreg_verbs.json")
|
||||
_KAIKKI = os.path.join(_HERE, "data", "kaikki_fr.jsonl")
|
||||
_CACHE = os.path.join(_HERE, "data", "fr_morph_cache.pkl")
|
||||
|
||||
# ── (mood, tense) -> UniMorph feature set that must ALL be present ────────────────
|
||||
_VERB_KEYMAP = {
|
||||
("ind", "present"): {"IND", "PRS"},
|
||||
("ind", "imperfect"): {"IND", "PST", "IPFV"}, # imparfait
|
||||
("ind", "passe_simple"): {"IND", "PST", "PFV"}, # passé simple
|
||||
("ind", "future"): {"IND", "FUT"},
|
||||
("ind", "conditional"): {"COND"}, # French: V;COND;1;SG
|
||||
("sbjv", "present"): {"SBJV", "PRS"},
|
||||
("sbjv", "imperfect"): {"SBJV", "PST"},
|
||||
("imp", "affirmative"): {"POS", "IMP"},
|
||||
}
|
||||
_PERSON = {"first": "1", "second": "2", "third": "3"}
|
||||
_NUMBER = {"singular": "SG", "plural": "PL"}
|
||||
|
||||
|
||||
def _feat_set(tag):
|
||||
return set(tag.split(";"))
|
||||
|
||||
|
||||
# ── build verb lexicon from UniMorph ─────────────────────────────────────────────
|
||||
def _build_verbs():
|
||||
verbs = {}
|
||||
part = {}
|
||||
ger = {}
|
||||
with open(_UNIMORPH, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
line = line.rstrip("\n")
|
||||
if not line or "\t" not in line:
|
||||
continue
|
||||
parts = line.split("\t")
|
||||
if len(parts) != 3:
|
||||
continue
|
||||
lemma, form, tag = parts
|
||||
f = _feat_set(tag)
|
||||
head = tag.split(";")[0]
|
||||
|
||||
if head == "V.PTCP":
|
||||
if "PST" in f:
|
||||
part.setdefault(lemma, form)
|
||||
elif "PRS" in f:
|
||||
ger.setdefault(lemma, form)
|
||||
continue
|
||||
if head == "V.CVB":
|
||||
if "PRS" in f:
|
||||
ger.setdefault(lemma, form)
|
||||
continue
|
||||
if head != "V":
|
||||
continue
|
||||
|
||||
person = next((p for p in ("1", "2", "3") if p in f), None)
|
||||
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
|
||||
if person is None or number is None:
|
||||
continue
|
||||
for (mood, tense), req in _VERB_KEYMAP.items():
|
||||
if not req <= f:
|
||||
continue
|
||||
if tense == "imperfect" and "PFV" in f:
|
||||
continue
|
||||
if tense == "passe_simple" and "IPFV" in f:
|
||||
continue
|
||||
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), form)
|
||||
break
|
||||
return verbs, part, ger
|
||||
|
||||
|
||||
# ── kaikki nouns + adjectives ────────────────────────────────────────────────────
|
||||
_EXCL_FORM_TAGS = {"alternative", "archaic", "obsolete", "dialectal", "regional",
|
||||
"diminutive", "augmentative", "pejorative", "comparative",
|
||||
"superlative", "misspelling", "rare", "informal", "literary",
|
||||
"poetic", "error-unrecognized-form", "construed", "collective",
|
||||
"nonstandard", "dated", "Louisiana", "Switzerland", "Belgium"}
|
||||
|
||||
|
||||
def _kaikki_gender(arg):
|
||||
if not arg:
|
||||
return None
|
||||
a = str(arg).lower()
|
||||
if a.startswith("f"):
|
||||
return "f"
|
||||
if a.startswith("m"):
|
||||
return "m"
|
||||
return None
|
||||
|
||||
|
||||
def _build_nouns_adjs():
|
||||
nouns = {}
|
||||
adjs = {}
|
||||
with open(_KAIKKI, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
try:
|
||||
d = json.loads(line)
|
||||
except Exception:
|
||||
continue
|
||||
pos = d.get("pos")
|
||||
word = d.get("word", "")
|
||||
if not word or " " in word:
|
||||
continue
|
||||
forms = d.get("forms", []) or []
|
||||
|
||||
if pos == "noun":
|
||||
ht = d.get("head_templates") or []
|
||||
g = None
|
||||
if ht:
|
||||
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
|
||||
if g is None:
|
||||
tags = d.get("tags") or []
|
||||
if "feminine" in tags:
|
||||
g = "f"
|
||||
elif "masculine" in tags:
|
||||
g = "m"
|
||||
pl = None
|
||||
for x in forms:
|
||||
t = set(x.get("tags") or [])
|
||||
if "plural" in t and not (t & _EXCL_FORM_TAGS):
|
||||
fm = x.get("form")
|
||||
if fm and " " not in fm and fm not in ("#", "-", "—"):
|
||||
pl = fm
|
||||
break
|
||||
if word not in nouns:
|
||||
nouns[word] = {"g": g, "SG": word, "PL": pl}
|
||||
else:
|
||||
cur = nouns[word]
|
||||
if cur.get("g") is None and g:
|
||||
cur["g"] = g
|
||||
if not cur.get("PL") and pl:
|
||||
cur["PL"] = pl
|
||||
|
||||
elif pos == "adj":
|
||||
d0 = adjs.setdefault(word, {})
|
||||
d0.setdefault(("m", "SG"), word)
|
||||
for x in forms:
|
||||
t = set(x.get("tags") or [])
|
||||
fm = x.get("form")
|
||||
if not fm or " " in fm or (t & _EXCL_FORM_TAGS):
|
||||
continue
|
||||
if "feminine" in t and "plural" in t:
|
||||
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
|
||||
elif "masculine" in t and "plural" in t:
|
||||
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
|
||||
elif "feminine" in t:
|
||||
d0[("f", "SG")] = d0.get(("f", "SG")) or fm
|
||||
elif "plural" in t:
|
||||
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
|
||||
return nouns, adjs
|
||||
|
||||
|
||||
def _build_cache():
|
||||
verbs, part, ger = _build_verbs()
|
||||
nouns, adjs = _build_nouns_adjs()
|
||||
with open(_IRREG, encoding="utf-8") as fh:
|
||||
irreg = json.load(fh)
|
||||
data = {"verbs": verbs, "part": part, "ger": ger,
|
||||
"nouns": nouns, "adjs": adjs, "irreg": irreg}
|
||||
try:
|
||||
with open(_CACHE, "wb") as fh:
|
||||
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
|
||||
except OSError:
|
||||
pass
|
||||
return data
|
||||
|
||||
|
||||
def _load():
|
||||
if os.path.exists(_CACHE):
|
||||
srcs = [_UNIMORPH, _KAIKKI, _IRREG]
|
||||
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
|
||||
if os.path.getmtime(_CACHE) >= newest:
|
||||
try:
|
||||
with open(_CACHE, "rb") as fh:
|
||||
return pickle.load(fh)
|
||||
except Exception:
|
||||
pass
|
||||
return _build_cache()
|
||||
|
||||
|
||||
_LEX = _load()
|
||||
_VERBS, _PART, _GER, _NOUNS, _ADJS, _IRREGV = (
|
||||
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"],
|
||||
_LEX["irreg"])
|
||||
|
||||
|
||||
# ── regular-ending rule fallback ─────────────────────────────────────────────────
|
||||
def _vclass(lemma):
|
||||
if lemma.endswith("er"):
|
||||
return "er"
|
||||
if lemma.endswith("ir"):
|
||||
return "ir"
|
||||
if lemma.endswith("re"):
|
||||
return "re"
|
||||
if lemma.endswith("oir"):
|
||||
return "oir"
|
||||
return None
|
||||
|
||||
|
||||
# present-tense endings [1sg,2sg,3sg,1pl,2pl,3pl]
|
||||
_REG_PRES = {
|
||||
"er": ["e", "es", "e", "ons", "ez", "ent"],
|
||||
"ir": ["is", "is", "it", "issons", "issez", "issent"], # -iss- class (finir)
|
||||
"re": ["s", "s", "", "ons", "ez", "ent"], # vendre: vends/vend
|
||||
}
|
||||
_REG_IMPF = ["ais", "ais", "ait", "ions", "iez", "aient"] # attaches to pres-1pl stem
|
||||
_REG_SUBJ = ["e", "es", "e", "ions", "iez", "ent"] # attaches to 3pl stem
|
||||
_REG_PS = { # passé simple
|
||||
"er": ["ai", "as", "a", "âmes", "âtes", "èrent"],
|
||||
"ir": ["is", "is", "it", "îmes", "îtes", "irent"],
|
||||
"re": ["is", "is", "it", "îmes", "îtes", "irent"],
|
||||
}
|
||||
_FUT = ["ai", "as", "a", "ons", "ez", "ont"]
|
||||
_COND = ["ais", "ais", "ait", "ions", "iez", "aient"]
|
||||
|
||||
|
||||
def _slot_idx(person, number):
|
||||
base = {"first": 0, "second": 1, "third": 2}[person]
|
||||
return base + (0 if number == "singular" else 3)
|
||||
|
||||
|
||||
def _fut_stem(lemma, vc):
|
||||
"""Future/conditional stem = infinitive (drop final -e of -re)."""
|
||||
if vc == "re":
|
||||
return lemma[:-1] # vendre -> vendr-
|
||||
return lemma # parler-, finir-
|
||||
|
||||
|
||||
def _pres_1pl_stem(lemma, vc):
|
||||
"""Imparfait stem = present 1pl minus -ons (parlons->parl-, finissons->finiss-)."""
|
||||
if vc == "er":
|
||||
stem = lemma[:-2]
|
||||
if stem.endswith("g"):
|
||||
return stem + "e" # mangeons -> mange- (imparfait mangeais)
|
||||
if stem.endswith("c"):
|
||||
return stem[:-1] + "ç" # commençons -> commenç-
|
||||
return stem
|
||||
if vc == "ir":
|
||||
return lemma[:-1] + "iss" # finir -> finiss-
|
||||
if vc == "re":
|
||||
return lemma[:-2] # vendre -> vend-
|
||||
return lemma[:-2]
|
||||
|
||||
|
||||
def _apply_er_spelling(stem, ending):
|
||||
"""-cer/-ger softening before a/o (commençons, mangeons)."""
|
||||
if ending and ending[0] in ("a", "o"):
|
||||
if stem.endswith("c"):
|
||||
return stem[:-1] + "ç" + ending
|
||||
if stem.endswith("g"):
|
||||
return stem + "e" + ending
|
||||
return stem + ending
|
||||
|
||||
|
||||
def _rule_conjugate(lemma, mood, tense, person, number):
|
||||
vc = _vclass(lemma)
|
||||
if vc is None:
|
||||
return None
|
||||
i = _slot_idx(person, number)
|
||||
|
||||
if mood == "ind" and tense in ("future", "conditional"):
|
||||
stem = _fut_stem(lemma, vc)
|
||||
end = (_FUT if tense == "future" else _COND)[i]
|
||||
return stem + end
|
||||
|
||||
if mood == "ind" and tense == "present":
|
||||
table = _REG_PRES.get("ir" if vc == "ir" else vc)
|
||||
if not table:
|
||||
return None
|
||||
body = lemma[:-2] if vc in ("er", "re") else lemma[:-1] if vc == "ir" else lemma[:-2]
|
||||
if vc == "ir":
|
||||
body = lemma[:-2] # fin- ; endings carry -iss-
|
||||
end = table[i]
|
||||
return body + end
|
||||
end = table[i]
|
||||
if vc == "er":
|
||||
return _apply_er_spelling(body, end)
|
||||
return body + end
|
||||
|
||||
if mood == "ind" and tense == "imperfect":
|
||||
stem = _pres_1pl_stem(lemma, vc)
|
||||
return stem + _REG_IMPF[i]
|
||||
|
||||
if mood == "ind" and tense == "passe_simple":
|
||||
table = _REG_PS.get("ir" if vc == "ir" else vc)
|
||||
if not table:
|
||||
return None
|
||||
body = lemma[:-2] if vc in ("er", "re") else lemma[:-2]
|
||||
end = table[i]
|
||||
if vc == "er":
|
||||
return _apply_er_spelling(body, end)
|
||||
return body + end
|
||||
|
||||
if mood == "sbjv" and tense == "present":
|
||||
# subjonctif: present-3pl stem + e/es/e/ions/iez/ent
|
||||
stem3 = _pres_1pl_stem(lemma, vc) if vc == "ir" else (
|
||||
lemma[:-2] if vc in ("er", "re") else lemma[:-2])
|
||||
if vc == "ir":
|
||||
stem3 = lemma[:-2] + "iss"
|
||||
end = _REG_SUBJ[i]
|
||||
if vc == "er":
|
||||
return _apply_er_spelling(stem3, end)
|
||||
return stem3 + end
|
||||
|
||||
if mood == "imp" and tense == "affirmative":
|
||||
# impératif ~ present indicative (tu drops -s for -er verbs)
|
||||
pres = _rule_conjugate(lemma, "ind", "present", person, number)
|
||||
if pres and vc == "er" and person == "second" and number == "singular":
|
||||
return pres[:-1] if pres.endswith("es") else pres
|
||||
return pres
|
||||
return None
|
||||
|
||||
|
||||
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
|
||||
def conjugate(lemma, mood, tense, person, number):
|
||||
"""Return (surface, confidence)."""
|
||||
lemma = lemma.strip().lower()
|
||||
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{number}"
|
||||
ir = _IRREGV.get(lemma)
|
||||
if ir and key in ir:
|
||||
return ir[key], "lexicon"
|
||||
p, n = _PERSON.get(person), _NUMBER.get(number)
|
||||
if p and n:
|
||||
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
r = _rule_conjugate(lemma, mood, tense, person, number)
|
||||
if r:
|
||||
return r, "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
# ── PUBLIC: participle + gerund/participe présent ────────────────────────────────
|
||||
def _participle_msg(lemma):
|
||||
ir = _IRREGV.get(lemma)
|
||||
if ir and "part" in ir:
|
||||
return ir["part"], "lexicon"
|
||||
if lemma in _PART:
|
||||
return _PART[lemma], "lexicon"
|
||||
return None, None
|
||||
|
||||
|
||||
# irregular participle fem/plural quirks (drop circonflexe: dû->due, dus)
|
||||
_PART_FIX = {"dû": {"f|SG": "due", "m|PL": "dus", "f|PL": "dues"}}
|
||||
|
||||
|
||||
def participle(lemma, gender="m", number="singular"):
|
||||
"""Past participle with French gender/number agreement.
|
||||
m.sg = base; f.sg = base+e; m.pl = base+s (invariable if base ends s/x);
|
||||
f.pl = f.sg+s."""
|
||||
lemma = lemma.strip().lower()
|
||||
g = "f" if gender == "f" else "m"
|
||||
num = "SG" if number == "singular" else "PL"
|
||||
msg, src = _participle_msg(lemma)
|
||||
conf = "lexicon"
|
||||
if msg is None:
|
||||
vc = _vclass(lemma)
|
||||
if vc == "er":
|
||||
msg = lemma[:-2] + "é"
|
||||
elif vc == "ir":
|
||||
msg = lemma[:-1] # finir -> fini, partir -> parti
|
||||
elif vc == "re":
|
||||
msg = lemma[:-2] + "u" # vendre -> vendu
|
||||
elif vc == "oir":
|
||||
msg = lemma[:-3] + "u" # (rough) recevoir handled by irreg
|
||||
else:
|
||||
return lemma, "fallback"
|
||||
conf = "rule"
|
||||
fix = _PART_FIX.get(msg)
|
||||
if fix and f"{g}|{num}" in fix:
|
||||
return fix[f"{g}|{num}"], conf
|
||||
if g == "m" and num == "SG":
|
||||
return msg, conf
|
||||
fem = msg + "e" if not msg.endswith("e") else msg
|
||||
if g == "f" and num == "SG":
|
||||
return fem, conf
|
||||
if g == "m" and num == "PL":
|
||||
return msg if msg.endswith(("s", "x")) else msg + "s", conf
|
||||
# f|PL
|
||||
return fem + "s", conf
|
||||
|
||||
|
||||
def gerund(lemma):
|
||||
"""Participe présent (base for gérondif 'en -ant')."""
|
||||
lemma = lemma.strip().lower()
|
||||
ir = _IRREGV.get(lemma)
|
||||
if ir and "ger" in ir:
|
||||
return ir["ger"], "lexicon"
|
||||
if lemma in _GER:
|
||||
return _GER[lemma], "lexicon"
|
||||
vc = _vclass(lemma)
|
||||
if vc == "er":
|
||||
stem = lemma[:-2]
|
||||
if stem.endswith("g"):
|
||||
return stem + "eant", "rule"
|
||||
if stem.endswith("c"):
|
||||
return stem[:-1] + "çant", "rule"
|
||||
return stem + "ant", "rule"
|
||||
if vc == "ir":
|
||||
return lemma[:-2] + "issant", "rule"
|
||||
if vc == "re":
|
||||
return lemma[:-2] + "ant", "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
|
||||
_FEM_SUF = ("tion", "sion", "aison", "ance", "ence", "ette", "elle", "esse",
|
||||
"ude", "ade", "ée", "té", "tié", "ie", "ise", "ure", "eur")
|
||||
_MASC_SUF = ("ment", "age", "eau", "isme", "oir", "ier", "eur", "in", "on")
|
||||
|
||||
|
||||
def _gender_heuristic(noun):
|
||||
for suf in _FEM_SUF:
|
||||
if noun.endswith(suf):
|
||||
return "f"
|
||||
for suf in _MASC_SUF:
|
||||
if noun.endswith(suf):
|
||||
return "m"
|
||||
if noun.endswith("e"):
|
||||
return "f"
|
||||
return "m"
|
||||
|
||||
|
||||
def noun_gender(lemma):
|
||||
lemma = lemma.strip().lower()
|
||||
d = _NOUNS.get(lemma)
|
||||
if d and d.get("g") in ("m", "f"):
|
||||
return d["g"]
|
||||
return _gender_heuristic(lemma)
|
||||
|
||||
|
||||
# closed sets for French plural irregularities
|
||||
_OU_X = {"bijou", "caillou", "chou", "genou", "hibou", "joujou", "pou"}
|
||||
_AIL_AUX = {"travail", "vitrail", "corail", "émail", "bail", "soupirail", "vantail"}
|
||||
_AL_S = {"bal", "carnaval", "festival", "récital", "chacal", "régal", "cal", "aval"}
|
||||
|
||||
|
||||
def _rule_plural(noun, gender):
|
||||
"""Deterministic French pluralization. (form, ok); ok=False FLAGS ambiguity."""
|
||||
if not noun:
|
||||
return noun, True
|
||||
if noun[-1:] in ("s", "x", "z"):
|
||||
return noun, True # invariable
|
||||
if noun in _OU_X:
|
||||
return noun + "x", True
|
||||
if noun.endswith(("eau", "au", "eu")):
|
||||
if noun in ("pneu", "bleu", "landau", "sarrau"):
|
||||
return noun + "s", True
|
||||
return noun + "x", True # bateau->bateaux, jeu->jeux
|
||||
if noun.endswith("al"):
|
||||
if noun in _AL_S:
|
||||
return noun + "s", True
|
||||
return noun[:-2] + "aux", True # cheval->chevaux
|
||||
if noun.endswith("ail"):
|
||||
if noun in _AIL_AUX:
|
||||
return noun[:-3] + "aux", True # travail->travaux
|
||||
return noun + "s", True
|
||||
return noun + "s", True # default
|
||||
|
||||
|
||||
def inflect_noun(lemma, number, gender=None):
|
||||
lemma = lemma.strip().lower()
|
||||
d = _NOUNS.get(lemma)
|
||||
if number == "singular":
|
||||
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
|
||||
if d and d.get("PL"):
|
||||
return d["PL"], "lexicon"
|
||||
g = gender or noun_gender(lemma)
|
||||
form, ok = _rule_plural(lemma, g)
|
||||
return form, ("rule" if ok else "fallback")
|
||||
|
||||
|
||||
# adjectives whose kaikki entries are unreliable: audited forms
|
||||
_ADJ_FIX = {
|
||||
"beau": {("m", "SG"): "beau", ("f", "SG"): "belle",
|
||||
("m", "PL"): "beaux", ("f", "PL"): "belles"},
|
||||
"nouveau": {("m", "SG"): "nouveau", ("f", "SG"): "nouvelle",
|
||||
("m", "PL"): "nouveaux", ("f", "PL"): "nouvelles"},
|
||||
"vieux": {("m", "SG"): "vieux", ("f", "SG"): "vieille",
|
||||
("m", "PL"): "vieux", ("f", "PL"): "vieilles"},
|
||||
"fou": {("m", "SG"): "fou", ("f", "SG"): "folle",
|
||||
("m", "PL"): "fous", ("f", "PL"): "folles"},
|
||||
"blanc": {("m", "SG"): "blanc", ("f", "SG"): "blanche",
|
||||
("m", "PL"): "blancs", ("f", "PL"): "blanches"},
|
||||
"long": {("m", "SG"): "long", ("f", "SG"): "longue",
|
||||
("m", "PL"): "longs", ("f", "PL"): "longues"},
|
||||
"bon": {("m", "SG"): "bon", ("f", "SG"): "bonne",
|
||||
("m", "PL"): "bons", ("f", "PL"): "bonnes"},
|
||||
}
|
||||
|
||||
|
||||
def _rule_fem(a):
|
||||
if a.endswith("e"):
|
||||
return a
|
||||
if a.endswith("er"):
|
||||
return a[:-2] + "ère"
|
||||
if a.endswith("eau"):
|
||||
return a[:-3] + "elle"
|
||||
if a.endswith("eux"):
|
||||
return a[:-3] + "euse"
|
||||
if a.endswith("f"):
|
||||
return a[:-1] + "ve"
|
||||
if a.endswith(("on", "en", "el", "eil", "et")):
|
||||
return a + a[-1] + "e" # bon->bonne, ancien->ancienne, muet->muette
|
||||
if a.endswith("c"):
|
||||
return a[:-1] + "che" # blanc->blanche (public->publique via FIX)
|
||||
return a + "e" # grand->grande, petit->petite, vert->verte
|
||||
|
||||
|
||||
def inflect_adj(lemma, gender, number):
|
||||
lemma = lemma.strip().lower()
|
||||
g = "f" if gender == "f" else "m"
|
||||
num = "SG" if number == "singular" else "PL"
|
||||
fix = _ADJ_FIX.get(lemma)
|
||||
if fix and (g, num) in fix:
|
||||
return fix[(g, num)], "lexicon"
|
||||
d = _ADJS.get(lemma)
|
||||
if d and d.get((g, num)):
|
||||
return d[(g, num)], "lexicon"
|
||||
# derive
|
||||
msc = (d.get(("m", "SG")) if d else None) or lemma
|
||||
if g == "m" and num == "SG":
|
||||
return msc, "lexicon" if d else "rule"
|
||||
fem = (d.get(("f", "SG")) if d else None) or _rule_fem(msc)
|
||||
if g == "f" and num == "SG":
|
||||
return fem, "lexicon" if (d and d.get(("f", "SG"))) else "rule"
|
||||
if g == "m" and num == "PL":
|
||||
if msc.endswith(("s", "x")):
|
||||
return msc, "rule"
|
||||
if msc.endswith("al"):
|
||||
return msc[:-2] + "aux", "rule"
|
||||
if msc.endswith("eau"):
|
||||
return msc + "x", "rule"
|
||||
return msc + "s", "rule"
|
||||
# f|PL
|
||||
return (fem if fem.endswith("s") else fem + "s"), "rule"
|
||||
|
||||
|
||||
def lexicon_stats():
|
||||
return {
|
||||
"verb_source": "UniMorph French (github.com/unimorph/fra) + kaikki.org "
|
||||
"irregulars (être + high-frequency)",
|
||||
"noun_adj_source": "kaikki.org French (Wiktionary extract)",
|
||||
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
|
||||
"unimorph_verb_forms": len(_VERBS),
|
||||
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
|
||||
"irregular_verb_lemmas": len(_IRREGV),
|
||||
"participle_lemmas": len(_PART),
|
||||
"gerund_lemmas": len(_GER),
|
||||
"noun_lemmas": len(_NOUNS),
|
||||
"adj_lemmas": len(_ADJS),
|
||||
}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
|
||||
tests = [
|
||||
("parler", "ind", "present", "first", "singular", "parle"),
|
||||
("être", "ind", "present", "third", "singular", "est"),
|
||||
("avoir", "ind", "present", "first", "singular", "ai"),
|
||||
("aller", "ind", "present", "third", "plural", "vont"),
|
||||
("finir", "ind", "present", "first", "singular", "finis"),
|
||||
("finir", "ind", "present", "first", "plural", "finissons"),
|
||||
("manger", "ind", "present", "first", "plural", "mangeons"),
|
||||
("faire", "ind", "future", "first", "singular", "ferai"),
|
||||
("pouvoir", "sbjv", "present", "third", "singular", "puisse"),
|
||||
("prendre", "ind", "passe_simple", "third", "singular", "prit"),
|
||||
("vendre", "ind", "present", "third", "singular", "vend"),
|
||||
("commencer", "ind", "imperfect", "first", "singular", "commençais"),
|
||||
]
|
||||
ok = 0
|
||||
for lemma, mood, tense, per, num, exp in tests:
|
||||
got, conf = conjugate(lemma, mood, tense, per, num)
|
||||
flag = "OK " if got == exp else "XX "
|
||||
ok += got == exp
|
||||
print(f" {flag}{lemma:10} {mood}/{tense:12} {per[:3]}.{num[:2]} -> {got:12} ({conf}) exp={exp}")
|
||||
print(f"verb tests {ok}/{len(tests)}")
|
||||
print(" gender: maison=", noun_gender("maison"), "chat=", noun_gender("chat"),
|
||||
"cheval=", noun_gender("cheval"), "nation=", noun_gender("nation"))
|
||||
print(" plural: cheval->", inflect_noun("cheval", "plural"),
|
||||
"| bateau->", inflect_noun("bateau", "plural"),
|
||||
"| prix->", inflect_noun("prix", "plural"),
|
||||
"| chat->", inflect_noun("chat", "plural"))
|
||||
print(" adj: petit/f/sg->", inflect_adj("petit", "f", "singular"),
|
||||
"| beau/f/sg->", inflect_adj("beau", "f", "singular"),
|
||||
"| heureux/f/sg->", inflect_adj("heureux", "f", "singular"),
|
||||
"| national/m/pl->", inflect_adj("national", "m", "plural"))
|
||||
print(" part: aller/f/sg->", participle("aller", "f", "singular"),
|
||||
"| prendre/f/pl->", participle("prendre", "f", "plural"),
|
||||
"| finir/m/pl->", participle("finir", "m", "plural"))
|
||||
print(" ger: manger->", gerund("manger"), "| finir->", gerund("finir"))
|
||||
@@ -0,0 +1,588 @@
|
||||
"""morphology_it_full.py — production-grade Italian morphological generator.
|
||||
|
||||
NOT a toy. Backed by three real, Wiktionary-lineage lexical sources:
|
||||
|
||||
VERBS
|
||||
UniMorph Italian (github.com/unimorph/ita, CC-BY-SA 3.0)
|
||||
10,009 verb lemmas × full paradigm, CLEAN orthography (no stress marks):
|
||||
indicative present / imperfetto (PST;IPFV) / passato remoto (PST;PFV) /
|
||||
futuro, condizionale (COND),
|
||||
congiuntivo presente (SBJV;PRS) / imperfetto (SBJV;PST),
|
||||
affirmative imperative, infinitive, gerundio (V.CVB;PRS),
|
||||
past participle (masc-sg; fem/plural derived by vowel rule).
|
||||
it_irreg_verbs.json — 66 high-frequency verbs UniMorph MISSES
|
||||
(essere, avere, potere, uscire, tenere, prendere, piacere, …), extracted
|
||||
from kaikki.org Italian, filtered to standard forms, and DE-STRESSED to
|
||||
real orthography (kaikki marks tonic stress everywhere: pàrlo->parlo,
|
||||
avùto->avuto; final legit accents kept: sarò, è). Built by build_it_irreg.py.
|
||||
This layer takes priority — it supplies the two auxiliaries essere/avere,
|
||||
which the whole passato-prossimo / essere-agreement system depends on.
|
||||
|
||||
NOUNS + ADJECTIVES — kaikki.org Italian (Wiktionary extract, CC-BY-SA 3.0)
|
||||
noun lemmas WITH inherent gender (head-template arg) + real (often irregular)
|
||||
plural — uomo->uomini, uovo->uova, dito->dita, città invariant — resolved
|
||||
PER LEMMA, never guessed.
|
||||
adjective lemmas with real feminine + masc/fem plural (italiano->italiana/
|
||||
italiani/italiane, felice->felici invariant).
|
||||
|
||||
Fallbacks (degrade, never crash, on OOV input):
|
||||
verbs : rule generator for regular -are/-ere/-ire (with -care/-gare h-insertion
|
||||
and -ciare/-giare/-iare i-drop spelling rules)
|
||||
nouns : gender heuristic (endings) + rule pluralization (ambiguous -co/-go FLAGGED)
|
||||
adjs : -o/-a/-e gender rule + rule pluralization
|
||||
|
||||
Confidence flag on every form:
|
||||
"lexicon" from UniMorph / kaikki-irregular / kaikki noun-adj (trust: high)
|
||||
"rule" deterministic rule (trust: medium)
|
||||
"fallback" could not inflect; returned lemma / ambiguous (trust: low -> FLAG)
|
||||
|
||||
Public API (used by realizer_it.py):
|
||||
conjugate(lemma, mood, tense, person, number) -> (form, conf)
|
||||
participle(lemma, gender="m", number="singular") -> (form, conf)
|
||||
gerund(lemma) -> (form, conf)
|
||||
noun_gender(lemma) -> "m"|"f"
|
||||
inflect_noun(lemma, number, gender=None) -> (form, conf)
|
||||
inflect_adj(lemma, gender, number) -> (form, conf)
|
||||
lexicon_stats() -> dict
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import pickle
|
||||
|
||||
_HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
_UNIMORPH = os.path.join(_HERE, "data", "ita.unimorph")
|
||||
_IRREG = os.path.join(_HERE, "data", "it_irreg_verbs.json")
|
||||
_KAIKKI = os.path.join(_HERE, "data", "kaikki_it.jsonl")
|
||||
_CACHE = os.path.join(_HERE, "data", "it_morph_cache.pkl")
|
||||
|
||||
# ── (mood, tense) -> UniMorph feature set that must ALL be present ────────────────
|
||||
_VERB_KEYMAP = {
|
||||
("ind", "present"): {"IND", "PRS"},
|
||||
("ind", "imperfect"): {"IND", "PST", "IPFV"},
|
||||
("ind", "passato_remoto"): {"IND", "PST", "PFV"},
|
||||
("ind", "future"): {"IND", "FUT"},
|
||||
("ind", "conditional"): {"COND"},
|
||||
("sbjv", "present"): {"SBJV", "PRS"},
|
||||
("sbjv", "imperfect"): {"SBJV", "PST"},
|
||||
("imp", "affirmative"): {"POS", "IMP"},
|
||||
}
|
||||
_PERSON = {"first": "1", "second": "2", "third": "3"}
|
||||
_NUMBER = {"singular": "SG", "plural": "PL"}
|
||||
|
||||
|
||||
def _feat_set(tag):
|
||||
return set(tag.split(";"))
|
||||
|
||||
|
||||
# ── build verb lexicon from UniMorph ─────────────────────────────────────────────
|
||||
def _build_verbs():
|
||||
verbs = {} # (lemma, "mood|tense|person|number") -> form
|
||||
part = {} # lemma -> masc-sg past participle
|
||||
ger = {} # lemma -> gerundio
|
||||
with open(_UNIMORPH, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
line = line.rstrip("\n")
|
||||
if not line or "\t" not in line:
|
||||
continue
|
||||
parts = line.split("\t")
|
||||
if len(parts) != 3:
|
||||
continue
|
||||
lemma, form, tag = parts
|
||||
f = _feat_set(tag)
|
||||
head = tag.split(";")[0]
|
||||
|
||||
if head == "V.PTCP":
|
||||
if "PST" in f:
|
||||
part.setdefault(lemma, form)
|
||||
continue
|
||||
if head == "V.CVB": # gerundio (converb, present)
|
||||
if "PRS" in f:
|
||||
ger.setdefault(lemma, form)
|
||||
continue
|
||||
if head != "V":
|
||||
continue
|
||||
|
||||
person = next((p for p in ("1", "2", "3") if p in f), None)
|
||||
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
|
||||
if person is None or number is None:
|
||||
continue
|
||||
for (mood, tense), req in _VERB_KEYMAP.items():
|
||||
# exact-set discipline: PST;PFV must not match PST;IPFV, etc.
|
||||
if not req <= f:
|
||||
continue
|
||||
# guard IND;PST ambiguity: require the specific aspect feature
|
||||
if tense == "imperfect" and "PFV" in f:
|
||||
continue
|
||||
if tense == "passato_remoto" and "IPFV" in f:
|
||||
continue
|
||||
# COND must not also be a subjunctive/imperative slot
|
||||
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), form)
|
||||
break
|
||||
return verbs, part, ger
|
||||
|
||||
|
||||
# ── kaikki nouns + adjectives ────────────────────────────────────────────────────
|
||||
_EXCL_FORM_TAGS = {"alternative", "archaic", "obsolete", "dialectal", "regional",
|
||||
"diminutive", "augmentative", "pejorative", "comparative",
|
||||
"superlative", "misspelling", "rare", "informal", "literary",
|
||||
"poetic", "error-unrecognized-form", "apocopic", "obsolete",
|
||||
"construed", "collective"}
|
||||
|
||||
|
||||
def _kaikki_gender(arg):
|
||||
if not arg:
|
||||
return None
|
||||
a = str(arg).lower()
|
||||
if a.startswith("f"):
|
||||
return "f"
|
||||
if a.startswith("m"):
|
||||
return "m"
|
||||
return None
|
||||
|
||||
|
||||
def _build_nouns_adjs():
|
||||
nouns = {} # lemma -> {"g","SG","PL"}
|
||||
adjs = {} # lemma -> {("m","SG"),("f","SG"),("m","PL"),("f","PL")}
|
||||
with open(_KAIKKI, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
try:
|
||||
d = json.loads(line)
|
||||
except Exception:
|
||||
continue
|
||||
pos = d.get("pos")
|
||||
word = d.get("word", "")
|
||||
if not word or " " in word:
|
||||
continue
|
||||
forms = d.get("forms", []) or []
|
||||
|
||||
if pos == "noun":
|
||||
ht = d.get("head_templates") or []
|
||||
g = None
|
||||
if ht:
|
||||
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
|
||||
if g is None:
|
||||
tags = d.get("tags") or []
|
||||
if "feminine" in tags:
|
||||
g = "f"
|
||||
elif "masculine" in tags:
|
||||
g = "m"
|
||||
pl = None
|
||||
for x in forms:
|
||||
t = set(x.get("tags") or [])
|
||||
if "plural" in t and not (t & _EXCL_FORM_TAGS):
|
||||
fm = x.get("form")
|
||||
if fm and " " not in fm and fm != "#":
|
||||
pl = fm
|
||||
break
|
||||
if word not in nouns:
|
||||
nouns[word] = {"g": g, "SG": word, "PL": pl}
|
||||
else:
|
||||
cur = nouns[word]
|
||||
if cur.get("g") is None and g:
|
||||
cur["g"] = g
|
||||
if not cur.get("PL") and pl:
|
||||
cur["PL"] = pl
|
||||
|
||||
elif pos == "adj":
|
||||
d0 = adjs.setdefault(word, {})
|
||||
d0.setdefault(("m", "SG"), word)
|
||||
for x in forms:
|
||||
t = set(x.get("tags") or [])
|
||||
fm = x.get("form")
|
||||
if not fm or " " in fm or (t & _EXCL_FORM_TAGS):
|
||||
continue
|
||||
if "feminine" in t and "plural" in t:
|
||||
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
|
||||
elif "masculine" in t and "plural" in t:
|
||||
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
|
||||
elif "feminine" in t:
|
||||
d0[("f", "SG")] = d0.get(("f", "SG")) or fm
|
||||
elif "plural" in t: # invariant-gender adj (felice -> felici)
|
||||
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
|
||||
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
|
||||
return nouns, adjs
|
||||
|
||||
|
||||
def _build_cache():
|
||||
verbs, part, ger = _build_verbs()
|
||||
nouns, adjs = _build_nouns_adjs()
|
||||
with open(_IRREG, encoding="utf-8") as fh:
|
||||
irreg = json.load(fh)
|
||||
data = {"verbs": verbs, "part": part, "ger": ger,
|
||||
"nouns": nouns, "adjs": adjs, "irreg": irreg}
|
||||
try:
|
||||
with open(_CACHE, "wb") as fh:
|
||||
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
|
||||
except OSError:
|
||||
pass
|
||||
return data
|
||||
|
||||
|
||||
def _load():
|
||||
if os.path.exists(_CACHE):
|
||||
srcs = [_UNIMORPH, _KAIKKI, _IRREG]
|
||||
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
|
||||
if os.path.getmtime(_CACHE) >= newest:
|
||||
try:
|
||||
with open(_CACHE, "rb") as fh:
|
||||
return pickle.load(fh)
|
||||
except Exception:
|
||||
pass
|
||||
return _build_cache()
|
||||
|
||||
|
||||
_LEX = _load()
|
||||
_VERBS, _PART, _GER, _NOUNS, _ADJS, _IRREGV = (
|
||||
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"],
|
||||
_LEX["irreg"])
|
||||
|
||||
|
||||
# ── regular-ending rule fallback ─────────────────────────────────────────────────
|
||||
def _vclass(lemma):
|
||||
if lemma.endswith("are"):
|
||||
return "are"
|
||||
if lemma.endswith("ere"):
|
||||
return "ere"
|
||||
if lemma.endswith("ire"):
|
||||
return "ire"
|
||||
return None
|
||||
|
||||
|
||||
# endings [1sg,2sg,3sg,1pl,2pl,3pl]
|
||||
_REG = {
|
||||
("ind", "present", "are"): ["o", "i", "a", "iamo", "ate", "ano"],
|
||||
("ind", "present", "ere"): ["o", "i", "e", "iamo", "ete", "ono"],
|
||||
("ind", "present", "ire"): ["o", "i", "e", "iamo", "ite", "ono"],
|
||||
("ind", "imperfect", "are"): ["avo", "avi", "ava", "avamo", "avate", "avano"],
|
||||
("ind", "imperfect", "ere"): ["evo", "evi", "eva", "evamo", "evate", "evano"],
|
||||
("ind", "imperfect", "ire"): ["ivo", "ivi", "iva", "ivamo", "ivate", "ivano"],
|
||||
("ind", "passato_remoto", "are"): ["ai", "asti", "ò", "ammo", "aste", "arono"],
|
||||
("ind", "passato_remoto", "ere"): ["ei", "esti", "é", "emmo", "este", "erono"],
|
||||
("ind", "passato_remoto", "ire"): ["ii", "isti", "ì", "immo", "iste", "irono"],
|
||||
("sbjv", "present", "are"): ["i", "i", "i", "iamo", "iate", "ino"],
|
||||
("sbjv", "present", "ere"): ["a", "a", "a", "iamo", "iate", "ano"],
|
||||
("sbjv", "present", "ire"): ["a", "a", "a", "iamo", "iate", "ano"],
|
||||
("sbjv", "imperfect", "are"): ["assi", "assi", "asse", "assimo", "aste", "assero"],
|
||||
("sbjv", "imperfect", "ere"): ["essi", "essi", "esse", "essimo", "este", "essero"],
|
||||
("sbjv", "imperfect", "ire"): ["issi", "issi", "isse", "issimo", "iste", "issero"],
|
||||
# imperative: 2sg,3sg(Lei),1pl,2pl,3pl (1sg has none)
|
||||
("imp", "affirmative", "are"): [None, "a", "i", "iamo", "ate", "ino"],
|
||||
("imp", "affirmative", "ere"): [None, "i", "a", "iamo", "ete", "ano"],
|
||||
("imp", "affirmative", "ire"): [None, "i", "a", "iamo", "ite", "ano"],
|
||||
}
|
||||
# future / conditional attach to a stem = infinitive minus final -e, with
|
||||
# -are -> -er (parlare->parler-), -ere/-ire keep (credere->creder-, dormir-)
|
||||
_FUT = ["ò", "ai", "à", "emo", "ete", "anno"]
|
||||
_COND = ["ei", "esti", "ebbe", "emmo", "este", "ebbero"]
|
||||
|
||||
|
||||
def _slot_idx(person, number):
|
||||
base = {"first": 0, "second": 1, "third": 2}[person]
|
||||
return base + (0 if number == "singular" else 3)
|
||||
|
||||
|
||||
def _fut_stem(lemma, vc):
|
||||
body = lemma[:-3] # drop are/ere/ire
|
||||
if vc == "are":
|
||||
return body + "er"
|
||||
return body + vc[0] + "r" # ere->er? no: keep vowel: creder-, dormir-
|
||||
# NOTE corrected below
|
||||
|
||||
|
||||
def _apply_are_spelling(stem, ending):
|
||||
"""-care/-gare insert h before front endings; -ciare/-giare/-sciare/-iare drop i."""
|
||||
front = ending[:1] in ("i", "e")
|
||||
if stem.endswith(("c", "g")) and front:
|
||||
return stem + "h" + ending
|
||||
if stem.endswith(("ci", "gi", "sci")) and ending[:1] == "i":
|
||||
return stem[:-1] + ending # mangi+iamo -> mangiamo
|
||||
if stem.endswith("i") and ending[:1] == "i":
|
||||
return stem[:-1] + ending # studi+iamo -> studiamo
|
||||
return stem + ending
|
||||
|
||||
|
||||
def _rule_conjugate(lemma, mood, tense, person, number):
|
||||
vc = _vclass(lemma)
|
||||
if vc is None:
|
||||
return None
|
||||
body = lemma[:-3]
|
||||
i = _slot_idx(person, number)
|
||||
if mood == "ind" and tense in ("future", "conditional"):
|
||||
stem = body + "er" if vc == "are" else body + vc[0] + "r"
|
||||
# ere: creder-, ire: dormir- -> body + 'e'/'i' + 'r'
|
||||
if vc == "ere":
|
||||
stem = body + "er"
|
||||
elif vc == "ire":
|
||||
stem = body + "ir"
|
||||
end = (_FUT if tense == "future" else _COND)[i]
|
||||
# spelling: -care/-gare -> cherò/gherò ; -ciare/-giare -> cerò/gerò
|
||||
if vc == "are":
|
||||
if body.endswith(("c", "g")):
|
||||
stem = body + "her"
|
||||
elif body.endswith(("ci", "gi", "sci")):
|
||||
stem = body[:-1] + "er"
|
||||
elif body.endswith("i"):
|
||||
stem = body[:-1] + "er"
|
||||
return stem + end
|
||||
table = _REG.get((mood, tense, vc))
|
||||
if not table:
|
||||
return None
|
||||
end = table[i]
|
||||
if end is None:
|
||||
return None
|
||||
if vc == "are":
|
||||
return _apply_are_spelling(body, end)
|
||||
# -ere/-ire: guard against double-i (dormi+iamo -> dormiamo)
|
||||
if body.endswith("i") and end[:1] == "i":
|
||||
return body[:-1] + end
|
||||
return body + end
|
||||
|
||||
|
||||
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
|
||||
def conjugate(lemma, mood, tense, person, number):
|
||||
"""Return (surface, confidence). mood in ind|sbjv|imp; tense per _VERB_KEYMAP."""
|
||||
lemma = lemma.strip().lower()
|
||||
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{number}"
|
||||
ir = _IRREGV.get(lemma)
|
||||
if ir and key in ir:
|
||||
return ir[key], "lexicon"
|
||||
p, n = _PERSON.get(person), _NUMBER.get(number)
|
||||
if p and n:
|
||||
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
r = _rule_conjugate(lemma, mood, tense, person, number)
|
||||
if r:
|
||||
return r, "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
# ── PUBLIC: participle + gerund ──────────────────────────────────────────────────
|
||||
def _participle_msg(lemma):
|
||||
"""Return (masc-sg participle, source) or (None, None)."""
|
||||
ir = _IRREGV.get(lemma)
|
||||
if ir and "part" in ir:
|
||||
return ir["part"], "lexicon"
|
||||
if lemma in _PART:
|
||||
return _PART[lemma], "lexicon"
|
||||
return None, None
|
||||
|
||||
|
||||
def participle(lemma, gender="m", number="singular"):
|
||||
"""Past participle with gender/number agreement (for essere-perfect & passives).
|
||||
UniMorph/irregular give masc-sg; fem/plural derived by final-vowel swap
|
||||
(-o -> -a/-i/-e), valid for regular -ato/-uto/-ito AND irregulars
|
||||
(preso->presa/presi/prese, aperto->aperta/aperti/aperte, morto->morta/...)."""
|
||||
lemma = lemma.strip().lower()
|
||||
g = "f" if gender == "f" else "m"
|
||||
num = "SG" if number == "singular" else "PL"
|
||||
msg, src = _participle_msg(lemma)
|
||||
conf = "lexicon"
|
||||
if msg is None:
|
||||
vc = _vclass(lemma)
|
||||
if vc == "are":
|
||||
msg = lemma[:-3] + "ato"
|
||||
elif vc == "ere":
|
||||
msg = lemma[:-3] + "uto"
|
||||
elif vc == "ire":
|
||||
msg = lemma[:-3] + "ito"
|
||||
else:
|
||||
return lemma, "fallback"
|
||||
conf = "rule"
|
||||
# agreement: only -o participles inflect for gender+number
|
||||
if msg.endswith("o"):
|
||||
stem = msg[:-1]
|
||||
suf = {"m|SG": "o", "f|SG": "a", "m|PL": "i", "f|PL": "e"}[f"{g}|{num}"]
|
||||
return stem + suf, conf
|
||||
return msg, conf # non -o participle: leave as-is (rare)
|
||||
|
||||
|
||||
def gerund(lemma):
|
||||
lemma = lemma.strip().lower()
|
||||
ir = _IRREGV.get(lemma)
|
||||
if ir and "ger" in ir:
|
||||
return ir["ger"], "lexicon"
|
||||
if lemma in _GER:
|
||||
return _GER[lemma], "lexicon"
|
||||
vc = _vclass(lemma)
|
||||
if vc == "are":
|
||||
return lemma[:-3] + "ando", "rule"
|
||||
if vc in ("ere", "ire"):
|
||||
return lemma[:-3] + "endo", "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
|
||||
_FEM_SUF = ("zione", "sione", "gione", "tà", "tù", "trice", "aggine", "udine",
|
||||
"igine", "ie", "essa", "izia", "ezza")
|
||||
_MASC_SUF = ("ore", "ame", "iere", "ale", "ile")
|
||||
|
||||
|
||||
def _gender_heuristic(noun):
|
||||
for suf in _FEM_SUF:
|
||||
if noun.endswith(suf):
|
||||
return "f"
|
||||
for suf in _MASC_SUF:
|
||||
if noun.endswith(suf):
|
||||
return "m"
|
||||
if noun.endswith("o"):
|
||||
return "m"
|
||||
if noun.endswith("a"):
|
||||
return "f"
|
||||
if noun.endswith("à") or noun.endswith("ù"):
|
||||
return "f"
|
||||
return "m" # -e and consonant-final loanwords default masculine
|
||||
|
||||
|
||||
def noun_gender(lemma):
|
||||
lemma = lemma.strip().lower()
|
||||
d = _NOUNS.get(lemma)
|
||||
if d and d.get("g") in ("m", "f"):
|
||||
return d["g"]
|
||||
return _gender_heuristic(lemma)
|
||||
|
||||
|
||||
def _rule_plural(noun, gender):
|
||||
"""Deterministic Italian pluralization. Returns (form, ok); ok=False FLAGS an
|
||||
ambiguous case the lexicon would normally resolve (-co/-go palatalization)."""
|
||||
if not noun:
|
||||
return noun, True
|
||||
# invariant: accented final vowel, consonant-final, monosyllable, -i final
|
||||
if noun[-1:] in ("à", "è", "é", "ì", "í", "ò", "ó", "ù", "ú"):
|
||||
return noun, True
|
||||
if noun[-1:] not in ("a", "e", "o", "i", "u"):
|
||||
return noun, True # consonant-final loanword: invariant
|
||||
if noun.endswith("i"):
|
||||
return noun, True # e.g. crisi, analisi: invariant
|
||||
if noun.endswith("io"):
|
||||
return noun[:-2] + "i", True # figlio->figli (unstressed i)
|
||||
if noun.endswith("cia") or noun.endswith("gia"):
|
||||
# vowel before cia/gia -> -cie/-gie ; consonant -> -ce/-ge (approx)
|
||||
return noun[:-2] + "e", True # arancia->arance (majority)
|
||||
if noun.endswith("ca"):
|
||||
return noun[:-2] + "che", True # amica->amiche
|
||||
if noun.endswith("ga"):
|
||||
return noun[:-2] + "ghe", True
|
||||
if noun.endswith("co"):
|
||||
return noun[:-2] + "chi", False # AMBIGUOUS (amico->amici) -> flag
|
||||
if noun.endswith("go"):
|
||||
return noun[:-2] + "ghi", False # AMBIGUOUS (psicologo->psicologi)
|
||||
if noun.endswith("a"):
|
||||
return noun[:-1] + "e", True # casa->case (m -a: -i, but rare)
|
||||
if noun.endswith("o"):
|
||||
return noun[:-1] + "i", True # libro->libri
|
||||
if noun.endswith("e"):
|
||||
return noun[:-1] + "i", True # cane->cani, chiave->chiavi
|
||||
return noun, True
|
||||
|
||||
|
||||
def inflect_noun(lemma, number, gender=None):
|
||||
lemma = lemma.strip().lower()
|
||||
d = _NOUNS.get(lemma)
|
||||
if number == "singular":
|
||||
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
|
||||
if d and d.get("PL"):
|
||||
return d["PL"], "lexicon"
|
||||
g = gender or noun_gender(lemma)
|
||||
form, ok = _rule_plural(lemma, g)
|
||||
return form, ("rule" if ok else "fallback")
|
||||
|
||||
|
||||
# adjectives whose kaikki entries are unreliable (messy inflection templates):
|
||||
# supply audited regular agreement forms (prenominal apocope handled in realizer).
|
||||
_ADJ_FIX = {
|
||||
"bello": {("m", "SG"): "bello", ("f", "SG"): "bella",
|
||||
("m", "PL"): "belli", ("f", "PL"): "belle"},
|
||||
"quello": {("m", "SG"): "quello", ("f", "SG"): "quella",
|
||||
("m", "PL"): "quelli", ("f", "PL"): "quelle"},
|
||||
}
|
||||
|
||||
|
||||
# ── PUBLIC: adjective agreement ──────────────────────────────────────────────────
|
||||
def inflect_adj(lemma, gender, number):
|
||||
lemma = lemma.strip().lower()
|
||||
g = "f" if gender == "f" else "m"
|
||||
num = "SG" if number == "singular" else "PL"
|
||||
fix = _ADJ_FIX.get(lemma)
|
||||
if fix and (g, num) in fix:
|
||||
return fix[(g, num)], "lexicon"
|
||||
d = _ADJS.get(lemma)
|
||||
if d:
|
||||
form = d.get((g, num))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
sg = d.get((g, "SG")) or d.get(("m", "SG")) or lemma
|
||||
if num == "PL":
|
||||
pl, ok = _rule_plural(sg, g)
|
||||
return pl, ("rule" if ok else "fallback")
|
||||
return sg, "lexicon"
|
||||
# rule fallback
|
||||
a = lemma
|
||||
if a.endswith("o"): # -o/-a/-i/-e class
|
||||
base = a[:-1]
|
||||
suf = {"m|SG": "o", "f|SG": "a", "m|PL": "i", "f|PL": "e"}[f"{g}|{num}"]
|
||||
return base + suf, "rule"
|
||||
if a.endswith("e"): # felice-class: SG invariant, PL -i
|
||||
if num == "PL":
|
||||
return a[:-1] + "i", "rule"
|
||||
return a, "rule"
|
||||
if num == "PL":
|
||||
p, ok = _rule_plural(a, g)
|
||||
return p, ("rule" if ok else "fallback")
|
||||
return a, "rule"
|
||||
|
||||
|
||||
def lexicon_stats():
|
||||
return {
|
||||
"verb_source": "UniMorph Italian (github.com/unimorph/ita) + kaikki.org "
|
||||
"irregulars (de-stressed)",
|
||||
"noun_adj_source": "kaikki.org Italian (Wiktionary extract)",
|
||||
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
|
||||
"unimorph_verb_forms": len(_VERBS),
|
||||
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
|
||||
"irregular_verb_lemmas": len(_IRREGV),
|
||||
"participle_lemmas": len(_PART),
|
||||
"gerund_lemmas": len(_GER),
|
||||
"noun_lemmas": len(_NOUNS),
|
||||
"adj_lemmas": len(_ADJS),
|
||||
}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
|
||||
tests = [
|
||||
("parlare", "ind", "present", "first", "singular", "parlo"),
|
||||
("essere", "ind", "present", "third", "singular", "è"),
|
||||
("avere", "ind", "present", "first", "singular", "ho"),
|
||||
("mangiare", "ind", "present", "second", "singular", "mangi"),
|
||||
("finire", "ind", "present", "first", "singular", "finisco"),
|
||||
("andare", "ind", "present", "third", "plural", "vanno"),
|
||||
("fare", "ind", "future", "first", "singular", "farò"),
|
||||
("potere", "sbjv", "present", "third", "singular", "possa"),
|
||||
("prendere", "ind", "passato_remoto", "first", "singular", "presi"),
|
||||
("cercare", "ind", "present", "second", "singular", "cerchi"),
|
||||
("dormire", "ind", "present", "third", "plural", "dormono"),
|
||||
("credere", "ind", "future", "first", "singular", "crederò"),
|
||||
]
|
||||
ok = 0
|
||||
for lemma, mood, tense, per, num, exp in tests:
|
||||
got, conf = conjugate(lemma, mood, tense, per, num)
|
||||
flag = "OK " if got == exp else "XX "
|
||||
ok += got == exp
|
||||
print(f" {flag}{lemma:9} {mood}/{tense:14} {per[:3]}.{num[:2]} -> {got:12} ({conf}) exp={exp}")
|
||||
print(f"verb tests {ok}/{len(tests)}")
|
||||
print(" gender: casa=", noun_gender("casa"), "problema=", noun_gender("problema"),
|
||||
"mano=", noun_gender("mano"), "città=", noun_gender("città"),
|
||||
"cane=", noun_gender("cane"))
|
||||
print(" plural: uomo->", inflect_noun("uomo", "plural"),
|
||||
"| uovo->", inflect_noun("uovo", "plural"),
|
||||
"| città->", inflect_noun("città", "plural"),
|
||||
"| amico->", inflect_noun("amico", "plural"),
|
||||
"| casa->", inflect_noun("casa", "plural"))
|
||||
print(" adj: italiano/f/pl->", inflect_adj("italiano", "f", "plural"),
|
||||
"| felice/m/pl->", inflect_adj("felice", "m", "plural"),
|
||||
"| bello/f/sg->", inflect_adj("bello", "f", "singular"))
|
||||
print(" part: aprire/f/sg->", participle("aprire", "f", "singular"),
|
||||
"| prendere/m/pl->", participle("prendere", "m", "plural"),
|
||||
"| andare/f/sg->", participle("andare", "f", "singular"))
|
||||
print(" ger: fare->", gerund("fare"), "| parlare->", gerund("parlare"))
|
||||
@@ -0,0 +1,666 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""morphology_lat_full.py — production-grade Latin morphological generator.
|
||||
|
||||
Latin is the FLAGSHIP dead-language realizer. It rides the *architecture* of the
|
||||
Romance/Italic engine (the same Realization / spec-driven design and the UniMorph
|
||||
loader pattern from morphology_it_full.py) but with the CASE SYSTEM RESTORED —
|
||||
the feature Romance lost. Latin therefore exercises machinery the modern Romance
|
||||
siblings never needed: 5 declensions x 6 cases x 2 numbers x 3 genders, plus a
|
||||
4-conjugation verb system with tense/mood/voice.
|
||||
|
||||
DATA (real, attested — no fabrication):
|
||||
|
||||
NOUNS + ADJECTIVES — UniMorph Latin (github.com/unimorph/lat, CC-BY-SA 3.0)
|
||||
163,182 N forms across ~thousands of lemmas, each with the full case paradigm
|
||||
N;NOM/GEN/DAT/ACC/ABL/VOC;SG/PL (real inflected forms, WITH macrons:
|
||||
puella->puellam, rēx->rēgis, corpus->corporis).
|
||||
244,197 ADJ forms with case x GENDER x number, incl. UniMorph's combined
|
||||
tags (GEN+DAT, MASC+FEM, MASC+FEM+NEUT) which are split on load.
|
||||
462,668 V.PTCP forms (participles) also carry case/gender/number.
|
||||
UniMorph N tags DO NOT encode inherent gender, so noun gender is inferred
|
||||
from the declension (nom-sg + gen-sg endings) with a curated exceptions
|
||||
map — the standard, attestable rule (1st decl -a/-ae = fem, 2nd -us/-i =
|
||||
masc, -um = neut, ...).
|
||||
|
||||
VERBS — RULE ENGINE (honest gap: UniMorph Latin's verb list is a 947-lemma
|
||||
sample of rare/prefixed verbs that MISSES every core textbook verb — amō,
|
||||
videō, sum, regō, ... are all absent). Latin conjugation is, however, highly
|
||||
regular, so verbs are generated by a deterministic 4-conjugation engine over
|
||||
curated principal parts (present / perfect / supine stems), sourced from
|
||||
standard references. Irregulars (sum, possum, eō, ferō, volō, nōlō, mālō)
|
||||
are curated full tables. Forms are flagged "rule" (not "lexicon") for honesty.
|
||||
|
||||
Confidence flag on every form (same contract as the Romance engine):
|
||||
"lexicon" from UniMorph (trust: high)
|
||||
"rule" deterministic morphology rule (trust: medium)
|
||||
"fallback" could not inflect; returned lemma (trust: low -> FLAG)
|
||||
|
||||
Public API (used by realizer_lat.py):
|
||||
decline_noun(lemma, case, number) -> (form, conf)
|
||||
noun_gender(lemma) -> "m"|"f"|"n"
|
||||
decline_adj(lemma, case, gender, number) -> (form, conf)
|
||||
conjugate(lemma, tense, mood, voice, person, number) -> (form, conf)
|
||||
participle(lemma, kind, case, gender, number) -> (form, conf) # kind: prs|pfv|fut
|
||||
infinitive(lemma, tense="present", voice="active") -> (form, conf)
|
||||
lexicon_stats() -> dict
|
||||
"""
|
||||
import os
|
||||
import pickle
|
||||
|
||||
_HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
_UNIMORPH = os.path.join(_HERE, "data", "lat.unimorph")
|
||||
_CACHE = os.path.join(_HERE, "data", "lat_morph_cache.pkl")
|
||||
|
||||
_CASES = ("NOM", "GEN", "DAT", "ACC", "ABL", "VOC")
|
||||
_CASE_MAP = {"nom": "NOM", "gen": "GEN", "dat": "DAT", "acc": "ACC",
|
||||
"abl": "ABL", "voc": "VOC"}
|
||||
_NUM = {"singular": "SG", "plural": "PL"}
|
||||
_GEN = {"m": "MASC", "f": "FEM", "n": "NEUT"}
|
||||
|
||||
|
||||
# ── UniMorph loader: noun + adjective + participle case paradigms ────────────────
|
||||
def _build_cache():
|
||||
nouns = {} # lemma -> {(CASE, NUM): form}
|
||||
adjs = {} # lemma -> {(CASE, GEN, NUM): form}
|
||||
ptcps = {} # lemma -> {(CASE, GEN, NUM): form} (from V.PTCP; keyed loosely)
|
||||
with open(_UNIMORPH, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
line = line.rstrip("\n")
|
||||
if not line or "\t" not in line:
|
||||
continue
|
||||
parts = line.split("\t")
|
||||
if len(parts) != 3:
|
||||
continue
|
||||
lemma, form, tag = parts
|
||||
feats = tag.split(";")
|
||||
head = feats[0]
|
||||
fs = set(feats)
|
||||
case = next((c for c in _CASES if c in fs), None)
|
||||
# handle combined case tags like GEN+DAT
|
||||
if case is None:
|
||||
for f in feats:
|
||||
if "+" in f and any(c in f.split("+") for c in _CASES):
|
||||
case = [c for c in _CASES if c in f.split("+")]
|
||||
break
|
||||
num = "SG" if "SG" in fs else ("PL" if "PL" in fs else None)
|
||||
if case is None or num is None:
|
||||
continue
|
||||
cases = case if isinstance(case, list) else [case]
|
||||
|
||||
if head == "N":
|
||||
d = nouns.setdefault(lemma, {})
|
||||
for c in cases:
|
||||
d.setdefault((c, num), form)
|
||||
elif head == "ADJ":
|
||||
# gender may be combined: MASC+FEM+NEUT, MASC+FEM
|
||||
genders = []
|
||||
for g in ("MASC", "FEM", "NEUT"):
|
||||
if any(g == x or (g in x.split("+")) for x in feats):
|
||||
genders.append(g)
|
||||
if not genders:
|
||||
genders = ["MASC", "FEM", "NEUT"]
|
||||
d = adjs.setdefault(lemma, {})
|
||||
for c in cases:
|
||||
for g in genders:
|
||||
d.setdefault((c, g, num), form)
|
||||
data = {"nouns": nouns, "adjs": adjs, "ptcps": ptcps}
|
||||
try:
|
||||
with open(_CACHE, "wb") as fh:
|
||||
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
|
||||
except OSError:
|
||||
pass
|
||||
return data
|
||||
|
||||
|
||||
def _load():
|
||||
if os.path.exists(_CACHE) and os.path.exists(_UNIMORPH):
|
||||
if os.path.getmtime(_CACHE) >= os.path.getmtime(_UNIMORPH):
|
||||
try:
|
||||
with open(_CACHE, "rb") as fh:
|
||||
return pickle.load(fh)
|
||||
except Exception:
|
||||
pass
|
||||
return _build_cache()
|
||||
|
||||
|
||||
_LEX = _load()
|
||||
_NOUNS, _ADJS = _LEX["nouns"], _LEX["adjs"]
|
||||
|
||||
|
||||
# ── noun gender inference (declension-based, curated exceptions) ─────────────────
|
||||
# Real, attestable rule: gender follows declension + nominative shape, with the
|
||||
# standard closed set of exceptions.
|
||||
_GENDER_EXC = {
|
||||
# 1st-declension masculines (people/agents)
|
||||
"agricola": "m", "poēta": "m", "nauta": "m", "incola": "m", "scrība": "m",
|
||||
"auriga": "m", "pīrāta": "m", "athlēta": "m",
|
||||
# 2nd-declension neuters / feminines
|
||||
"vīrus": "n", "vulgus": "n", "pelagus": "n", "humus": "f",
|
||||
# common 3rd-declension whose gender the ending would mispredict
|
||||
"rēx": "m", "dux": "m", "mīles": "m", "pater": "m", "frāter": "m",
|
||||
"homō": "m", "leō": "m", "sōl": "m", "mōns": "m", "pōns": "m", "fōns": "m",
|
||||
"sanguis": "m", "ōrdō": "m", "sermō": "m", "amor": "m", "dolor": "m",
|
||||
"labor": "m", "timor": "m", "honor": "m", "color": "m", "pēs": "m",
|
||||
"dēns": "m", "flōs": "m", "mōs": "m", "mensis": "m", "orbis": "m",
|
||||
"piscis": "m", "ignis": "m", "collis": "m", "grex": "m", "prīnceps": "m",
|
||||
"māter": "f", "soror": "f", "uxor": "f", "mulier": "f", "virgō": "f",
|
||||
"urbs": "f", "arx": "f", "pāx": "f", "lēx": "f", "lūx": "f", "vōx": "f",
|
||||
"nox": "f", "nix": "f", "vīs": "f", "salūs": "f", "virtūs": "f",
|
||||
"aetās": "f", "cīvitās": "f", "lībertās": "f", "vēritās": "f", "voluptās": "f",
|
||||
"nātiō": "f", "ratiō": "f", "ōrātiō": "f", "legiō": "f", "regiō": "f",
|
||||
"mens": "f", "gens": "f", "ars": "f", "pars": "f", "mors": "f", "sors": "f",
|
||||
"nāvis": "f", "turris": "f", "avis": "f", "vallis": "f", "classis": "f",
|
||||
"corpus": "n", "tempus": "n", "opus": "n", "genus": "n", "onus": "n",
|
||||
"pectus": "n", "latus": "n", "vulnus": "n", "scelus": "n", "sīdus": "n",
|
||||
"caput": "n", "iter": "n", "flūmen": "n", "nōmen": "n", "carmen": "n",
|
||||
"agmen": "n", "certāmen": "n", "lūmen": "n", "ōmen": "n", "cōgnōmen": "n",
|
||||
"mare": "n", "animal": "n", "exemplar": "n", "rēte": "n",
|
||||
# 4th-declension exceptions
|
||||
"manus": "f", "domus": "f", "tribus": "f", "porticus": "f", "īdūs": "f",
|
||||
"cornū": "n", "genū": "n", "gelū": "n", "verū": "n",
|
||||
# 5th-declension
|
||||
"diēs": "m", "merīdiēs": "m",
|
||||
}
|
||||
|
||||
|
||||
def _infer_gender(lemma):
|
||||
if lemma in _GENDER_EXC:
|
||||
return _GENDER_EXC[lemma]
|
||||
d = _NOUNS.get(lemma)
|
||||
nom = d.get(("NOM", "SG")) if d else lemma
|
||||
gen = d.get(("GEN", "SG")) if d else None
|
||||
nom = nom or lemma
|
||||
# 5th declension: gen -eī / -ēī
|
||||
if gen and (gen.endswith("eī") or gen.endswith("ēī")):
|
||||
return "f"
|
||||
# 1st declension: nom -a, gen -ae
|
||||
if nom.endswith("a") and (not gen or gen.endswith("ae")):
|
||||
return "f"
|
||||
# 2nd declension neuter: nom -um
|
||||
if nom.endswith("um"):
|
||||
return "n"
|
||||
# 2nd declension masc: nom -us/-er/-ir, gen -ī
|
||||
if (nom.endswith("us") or nom.endswith("er") or nom.endswith("ir")) and \
|
||||
(not gen or gen.endswith("ī")):
|
||||
return "m"
|
||||
# 4th declension: gen -ūs
|
||||
if gen and gen.endswith("ūs"):
|
||||
return "n" if nom.endswith("ū") else "m"
|
||||
# 3rd declension neuters by common nom endings
|
||||
if nom.endswith(("men", "us", "ur", "al", "ar", "e", "ma")):
|
||||
# -us here is 3rd-decl neuter type (corpus) only if gen shows -oris/-eris
|
||||
if nom.endswith("us") and gen and (gen.endswith("oris") or gen.endswith("eris")
|
||||
or gen.endswith("uris")):
|
||||
return "n"
|
||||
if nom.endswith(("men", "al", "ar", "e")):
|
||||
return "n"
|
||||
# default 3rd-declension: masculine (most common)
|
||||
return "m"
|
||||
|
||||
|
||||
_GENDER_CACHE = {}
|
||||
|
||||
|
||||
def noun_gender(lemma):
|
||||
lemma = lemma.strip()
|
||||
if lemma not in _GENDER_CACHE:
|
||||
_GENDER_CACHE[lemma] = _infer_gender(lemma)
|
||||
return _GENDER_CACHE[lemma]
|
||||
|
||||
|
||||
# ── PUBLIC: noun declension ─────────────────────────────────────────────────────
|
||||
def decline_noun(lemma, case, number):
|
||||
lemma = lemma.strip()
|
||||
C = _CASE_MAP.get(case, case.upper())
|
||||
N = _NUM.get(number, number)
|
||||
d = _NOUNS.get(lemma)
|
||||
if d and (C, N) in d:
|
||||
return d[(C, N)], "lexicon"
|
||||
# abl sg often == the -e/-o form; try nom fallback
|
||||
if d:
|
||||
# try VOC==NOM, ACC neuter==NOM etc are already in data; last resort lemma
|
||||
return lemma, "fallback"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
# ── PUBLIC: adjective declension ────────────────────────────────────────────────
|
||||
def decline_adj(lemma, case, gender, number):
|
||||
lemma = lemma.strip()
|
||||
C = _CASE_MAP.get(case, case.upper())
|
||||
G = _GEN.get(gender, gender.upper())
|
||||
N = _NUM.get(number, number)
|
||||
d = _ADJS.get(lemma)
|
||||
if d and (C, G, N) in d:
|
||||
return d[(C, G, N)], "lexicon"
|
||||
# try other gender (some adjs listed only under MASC+FEM etc handled at load)
|
||||
if d:
|
||||
for altG in ("MASC", "FEM", "NEUT"):
|
||||
if (C, altG, N) in d:
|
||||
return d[(C, altG, N)], "lexicon"
|
||||
return lemma, "fallback"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
# ═══════════════════════════════════════════════════════════════════════════════
|
||||
# VERB RULE ENGINE (4 conjugations + curated irregulars)
|
||||
# ═══════════════════════════════════════════════════════════════════════════════
|
||||
# Curated principal parts for common attested verbs:
|
||||
# lemma -> (conj, present_stem, perfect_stem, supine_stem)
|
||||
# conj in {1,2,3,"3io",4}. Stems carry macrons (matching UniMorph orthography).
|
||||
_VERBS = {
|
||||
"amō": (1, "am", "amāv", "amāt"),
|
||||
"laudō": (1, "laud", "laudāv", "laudāt"),
|
||||
"portō": (1, "port", "portāv", "portāt"),
|
||||
"vocō": (1, "voc", "vocāv", "vocāt"),
|
||||
"dō": (1, "d", "ded", "dat"),
|
||||
"spectō": (1, "spect", "spectāv", "spectāt"),
|
||||
"pugnō": (1, "pugn", "pugnāv", "pugnāt"),
|
||||
"labōrō": (1, "labōr", "labōrāv", "labōrāt"),
|
||||
"necō": (1, "nec", "necāv", "necāt"),
|
||||
"parō": (1, "par", "parāv", "parāt"),
|
||||
"cōgitō": (1, "cōgit", "cōgitāv", "cōgitāt"),
|
||||
"habitō": (1, "habit", "habitāv", "habitāt"),
|
||||
"nārrō": (1, "nārr", "nārrāv", "nārrāt"),
|
||||
"servō": (1, "serv", "servāv", "servāt"),
|
||||
"superō": (1, "super", "superāv", "superāt"),
|
||||
"oppugnō": (1, "oppugn", "oppugnāv", "oppugnāt"),
|
||||
"ambulō": (1, "ambul", "ambulāv", "ambulāt"),
|
||||
"clāmō": (1, "clām", "clāmāv", "clāmāt"),
|
||||
"vulnerō": (1, "vulner", "vulnerāv", "vulnerāt"),
|
||||
"aedificō": (1, "aedific", "aedificāv", "aedificāt"),
|
||||
"expugnō": (1, "expugn", "expugnāv", "expugnāt"),
|
||||
"dēfendō": (3, "dēfend", "dēfend", "dēfēns"),
|
||||
"petō": (3, "pet", "petīv", "petīt"),
|
||||
"occīdō": (3, "occīd", "occīd", "occīs"),
|
||||
"interficiō": ("3io", "interfic", "interfēc", "interfect"),
|
||||
"timeō": (2, "tim", "timu", None),
|
||||
"iaceō": (2, "iac", "iacu", None),
|
||||
"pāreō": (2, "pār", "pāru", "pārit"),
|
||||
"respondeō": (2, "respond", "respond", "respōns"),
|
||||
"vertō": (3, "vert", "vert", "vers"),
|
||||
"ostendō": (3, "ostend", "ostend", "ostent"),
|
||||
"cōnstituō": (3, "cōnstitu", "cōnstitu", "cōnstitūt"),
|
||||
"cōgnōscō": (3, "cōgnōsc", "cōgnōv", "cōgnit"),
|
||||
"crēdō": (3, "crēd", "crēdid", "crēdit"),
|
||||
"ēdūcō": (3, "ēdūc", "ēdūx", "ēduct"),
|
||||
"cōnservō": (1, "cōnserv", "cōnservāv", "cōnservāt"),
|
||||
"iuvō": (1, "iuv", "iūv", "iūt"),
|
||||
"dēbeō": (2, "dēb", "dēbu", "dēbit"),
|
||||
"moneō": (2, "mon", "monu", "monit"),
|
||||
"videō": (2, "vid", "vīd", "vīs"),
|
||||
"habeō": (2, "hab", "habu", "habit"),
|
||||
"teneō": (2, "ten", "tenu", "tent"),
|
||||
"timeō": (2, "tim", "timu", None),
|
||||
"terreō": (2, "terr", "terru", "territ"),
|
||||
"dēleō": (2, "dēl", "dēlēv", "dēlēt"),
|
||||
"iubeō": (2, "iub", "iuss", "iuss"),
|
||||
"maneō": (2, "man", "māns", "māns"),
|
||||
"moveō": (2, "mov", "mōv", "mōt"),
|
||||
"doceō": (2, "doc", "docu", "doct"),
|
||||
"sedeō": (2, "sed", "sēd", "sess"),
|
||||
"rīdeō": (2, "rīd", "rīs", "rīs"),
|
||||
"regō": (3, "reg", "rēx", "rēct"),
|
||||
"dūcō": (3, "dūc", "dūx", "duct"),
|
||||
"scrībō": (3, "scrīb", "scrīps", "scrīpt"),
|
||||
"mittō": (3, "mitt", "mīs", "miss"),
|
||||
"pōnō": (3, "pōn", "posu", "posit"),
|
||||
"agō": (3, "ag", "ēg", "āct"),
|
||||
"dīcō": (3, "dīc", "dīx", "dict"),
|
||||
"gerō": (3, "ger", "gess", "gest"),
|
||||
"vincō": (3, "vinc", "vīc", "vict"),
|
||||
"petō": (3, "pet", "petīv", "petīt"),
|
||||
"legō": (3, "leg", "lēg", "lēct"),
|
||||
"currō": (3, "curr", "cucurr", "curs"),
|
||||
"vīvō": (3, "vīv", "vīx", "vīct"),
|
||||
"quaerō": (3, "quaer", "quaesīv", "quaesīt"),
|
||||
"trahō": (3, "trah", "trāx", "tract"),
|
||||
"claudō": (3, "claud", "claus", "claus"),
|
||||
"cōgō": (3, "cōg", "coēg", "coāct"),
|
||||
"relinquō": (3, "relinqu", "relīqu", "relict"),
|
||||
"capiō": ("3io", "cap", "cēp", "capt"),
|
||||
"faciō": ("3io", "fac", "fēc", "fact"),
|
||||
"iaciō": ("3io", "iac", "iēc", "iact"),
|
||||
"rapiō": ("3io", "rap", "rapu", "rapt"),
|
||||
"fugiō": ("3io", "fug", "fūg", "fugit"),
|
||||
"cupiō": ("3io", "cup", "cupīv", "cupīt"),
|
||||
"accipiō": ("3io", "accip", "accēp", "accept"),
|
||||
"audiō": (4, "aud", "audīv", "audīt"),
|
||||
"veniō": (4, "ven", "vēn", "vent"),
|
||||
"sciō": (4, "sc", "scīv", "scīt"),
|
||||
"sentiō": (4, "sent", "sēns", "sēns"),
|
||||
"mūniō": (4, "mūn", "mūnīv", "mūnīt"),
|
||||
"dormiō": (4, "dorm", "dormīv", "dormīt"),
|
||||
"aperiō": (4, "aper", "aperu", "apert"),
|
||||
"inveniō": (4, "inven", "invēn", "invent"),
|
||||
}
|
||||
|
||||
# ── Present-system paradigms: full ending tables per conjugation, attached to the
|
||||
# bare present stem (pstem). Hardcoded from the standard grammar with correct
|
||||
# macrons/vowel-lengths — deterministic and independently verifiable. Keys:
|
||||
# (tense, mood, voice) -> {conj: [1sg,2sg,3sg,1pl,2pl,3pl]}
|
||||
_PARADIGM = {
|
||||
("present", "ind", "active"): {
|
||||
1: ["ō", "ās", "at", "āmus", "ātis", "ant"],
|
||||
2: ["eō", "ēs", "et", "ēmus", "ētis", "ent"],
|
||||
3: ["ō", "is", "it", "imus", "itis", "unt"],
|
||||
"3io": ["iō", "is", "it", "imus", "itis", "iunt"],
|
||||
4: ["iō", "īs", "it", "īmus", "ītis", "iunt"],
|
||||
},
|
||||
("present", "ind", "passive"): {
|
||||
1: ["or", "āris", "ātur", "āmur", "āminī", "antur"],
|
||||
2: ["eor", "ēris", "ētur", "ēmur", "ēminī", "entur"],
|
||||
3: ["or", "eris", "itur", "imur", "iminī", "untur"],
|
||||
"3io": ["ior", "eris", "itur", "imur", "iminī", "iuntur"],
|
||||
4: ["ior", "īris", "ītur", "īmur", "īminī", "iuntur"],
|
||||
},
|
||||
("imperfect", "ind", "active"): {
|
||||
1: ["ābam", "ābās", "ābat", "ābāmus", "ābātis", "ābant"],
|
||||
2: ["ēbam", "ēbās", "ēbat", "ēbāmus", "ēbātis", "ēbant"],
|
||||
3: ["ēbam", "ēbās", "ēbat", "ēbāmus", "ēbātis", "ēbant"],
|
||||
"3io": ["iēbam", "iēbās", "iēbat", "iēbāmus", "iēbātis", "iēbant"],
|
||||
4: ["iēbam", "iēbās", "iēbat", "iēbāmus", "iēbātis", "iēbant"],
|
||||
},
|
||||
("imperfect", "ind", "passive"): {
|
||||
1: ["ābar", "ābāris", "ābātur", "ābāmur", "ābāminī", "ābantur"],
|
||||
2: ["ēbar", "ēbāris", "ēbātur", "ēbāmur", "ēbāminī", "ēbantur"],
|
||||
3: ["ēbar", "ēbāris", "ēbātur", "ēbāmur", "ēbāminī", "ēbantur"],
|
||||
"3io": ["iēbar", "iēbāris", "iēbātur", "iēbāmur", "iēbāminī", "iēbantur"],
|
||||
4: ["iēbar", "iēbāris", "iēbātur", "iēbāmur", "iēbāminī", "iēbantur"],
|
||||
},
|
||||
("future", "ind", "active"): {
|
||||
1: ["ābō", "ābis", "ābit", "ābimus", "ābitis", "ābunt"],
|
||||
2: ["ēbō", "ēbis", "ēbit", "ēbimus", "ēbitis", "ēbunt"],
|
||||
3: ["am", "ēs", "et", "ēmus", "ētis", "ent"],
|
||||
"3io": ["iam", "iēs", "iet", "iēmus", "iētis", "ient"],
|
||||
4: ["iam", "iēs", "iet", "iēmus", "iētis", "ient"],
|
||||
},
|
||||
("future", "ind", "passive"): {
|
||||
1: ["ābor", "āberis", "ābitur", "ābimur", "ābiminī", "ābuntur"],
|
||||
2: ["ēbor", "ēberis", "ēbitur", "ēbimur", "ēbiminī", "ēbuntur"],
|
||||
3: ["ar", "ēris", "ētur", "ēmur", "ēminī", "entur"],
|
||||
"3io": ["iar", "iēris", "iētur", "iēmur", "iēminī", "ientur"],
|
||||
4: ["iar", "iēris", "iētur", "iēmur", "iēminī", "ientur"],
|
||||
},
|
||||
("present", "sbjv", "active"): {
|
||||
1: ["em", "ēs", "et", "ēmus", "ētis", "ent"],
|
||||
2: ["eam", "eās", "eat", "eāmus", "eātis", "eant"],
|
||||
3: ["am", "ās", "at", "āmus", "ātis", "ant"],
|
||||
"3io": ["iam", "iās", "iat", "iāmus", "iātis", "iant"],
|
||||
4: ["iam", "iās", "iat", "iāmus", "iātis", "iant"],
|
||||
},
|
||||
("present", "sbjv", "passive"): {
|
||||
1: ["er", "ēris", "ētur", "ēmur", "ēminī", "entur"],
|
||||
2: ["ear", "eāris", "eātur", "eāmur", "eāminī", "eantur"],
|
||||
3: ["ar", "āris", "ātur", "āmur", "āminī", "antur"],
|
||||
"3io": ["iar", "iāris", "iātur", "iāmur", "iāminī", "iantur"],
|
||||
4: ["iar", "iāris", "iātur", "iāmur", "iāminī", "iantur"],
|
||||
},
|
||||
("imperfect", "sbjv", "active"): {
|
||||
1: ["ārem", "ārēs", "āret", "ārēmus", "ārētis", "ārent"],
|
||||
2: ["ērem", "ērēs", "ēret", "ērēmus", "ērētis", "ērent"],
|
||||
3: ["erem", "erēs", "eret", "erēmus", "erētis", "erent"],
|
||||
"3io": ["erem", "erēs", "eret", "erēmus", "erētis", "erent"],
|
||||
4: ["īrem", "īrēs", "īret", "īrēmus", "īrētis", "īrent"],
|
||||
},
|
||||
("imperfect", "sbjv", "passive"): {
|
||||
1: ["ārer", "ārēris", "ārētur", "ārēmur", "ārēminī", "ārentur"],
|
||||
2: ["ērer", "ērēris", "ērētur", "ērēmur", "ērēminī", "ērentur"],
|
||||
3: ["erer", "erēris", "erētur", "erēmur", "erēminī", "erentur"],
|
||||
"3io": ["erer", "erēris", "erētur", "erēmur", "erēminī", "erentur"],
|
||||
4: ["īrer", "īrēris", "īrētur", "īrēmur", "īrēminī", "īrentur"],
|
||||
},
|
||||
}
|
||||
# perfect-active endings (added to perfect stem) — same for all conjugations
|
||||
_PERF_ACT = {
|
||||
("perfect", "ind"): ["ī", "istī", "it", "imus", "istis", "ērunt"],
|
||||
("pluperfect", "ind"): ["eram", "erās", "erat", "erāmus", "erātis", "erant"],
|
||||
("futureperfect", "ind"): ["erō", "eris", "erit", "erimus", "eritis", "erint"],
|
||||
("perfect", "sbjv"): ["erim", "erīs", "erit", "erīmus", "erītis", "erint"],
|
||||
("pluperfect", "sbjv"):["issem", "issēs", "isset", "issēmus", "issētis", "issent"],
|
||||
}
|
||||
|
||||
|
||||
def _idx(person, number):
|
||||
base = {"first": 0, "second": 1, "third": 2}[person]
|
||||
return base + (0 if number == "singular" else 3)
|
||||
|
||||
|
||||
def _present_system(conj, pstem, tense, mood, voice, person, number):
|
||||
"""Generate a present-system form (present/imperfect/future ind & subj)."""
|
||||
table = _PARADIGM.get((tense, mood, voice))
|
||||
if not table or conj not in table:
|
||||
return None
|
||||
return pstem + table[conj][_idx(person, number)]
|
||||
|
||||
|
||||
def _active_infinitive_stem(conj, pstem):
|
||||
return {1: pstem + "ā", 2: pstem + "ē", 3: pstem + "e",
|
||||
"3io": pstem + "e", 4: pstem + "ī"}[conj]
|
||||
|
||||
|
||||
_IRREG = {
|
||||
"sum": {
|
||||
("present", "ind", "active"): ["sum", "es", "est", "sumus", "estis", "sunt"],
|
||||
("imperfect", "ind", "active"): ["eram", "erās", "erat", "erāmus", "erātis", "erant"],
|
||||
("future", "ind", "active"): ["erō", "eris", "erit", "erimus", "eritis", "erunt"],
|
||||
("perfect", "ind", "active"): ["fuī", "fuistī", "fuit", "fuimus", "fuistis", "fuērunt"],
|
||||
("pluperfect", "ind", "active"): ["fueram", "fuerās", "fuerat", "fuerāmus", "fuerātis", "fuerant"],
|
||||
("present", "sbjv", "active"): ["sim", "sīs", "sit", "sīmus", "sītis", "sint"],
|
||||
("imperfect", "sbjv", "active"): ["essem", "essēs", "esset", "essēmus", "essētis", "essent"],
|
||||
},
|
||||
"possum": {
|
||||
("present", "ind", "active"): ["possum", "potes", "potest", "possumus", "potestis", "possunt"],
|
||||
("imperfect", "ind", "active"): ["poteram", "poterās", "poterat", "poterāmus", "poterātis", "poterant"],
|
||||
("future", "ind", "active"): ["poterō", "poteris", "poterit", "poterimus", "poteritis", "poterunt"],
|
||||
("perfect", "ind", "active"): ["potuī", "potuistī", "potuit", "potuimus", "potuistis", "potuērunt"],
|
||||
("present", "sbjv", "active"): ["possim", "possīs", "possit", "possīmus", "possītis", "possint"],
|
||||
},
|
||||
"eō": {
|
||||
("present", "ind", "active"): ["eō", "īs", "it", "īmus", "ītis", "eunt"],
|
||||
("imperfect", "ind", "active"): ["ībam", "ībās", "ībat", "ībāmus", "ībātis", "ībant"],
|
||||
("future", "ind", "active"): ["ībō", "ībis", "ībit", "ībimus", "ībitis", "ībunt"],
|
||||
("perfect", "ind", "active"): ["iī", "īstī", "iit", "iimus", "īstis", "iērunt"],
|
||||
("present", "sbjv", "active"): ["eam", "eās", "eat", "eāmus", "eātis", "eant"],
|
||||
},
|
||||
"volō": {
|
||||
("present", "ind", "active"): ["volō", "vīs", "vult", "volumus", "vultis", "volunt"],
|
||||
("imperfect", "ind", "active"): ["volēbam", "volēbās", "volēbat", "volēbāmus", "volēbātis", "volēbant"],
|
||||
("future", "ind", "active"): ["volam", "volēs", "volet", "volēmus", "volētis", "volent"],
|
||||
("perfect", "ind", "active"): ["voluī", "voluistī", "voluit", "voluimus", "voluistis", "voluērunt"],
|
||||
("present", "sbjv", "active"): ["velim", "velīs", "velit", "velīmus", "velītis", "velint"],
|
||||
},
|
||||
"nōlō": {
|
||||
("present", "ind", "active"): ["nōlō", "nōn vīs", "nōn vult", "nōlumus", "nōn vultis", "nōlunt"],
|
||||
("present", "sbjv", "active"): ["nōlim", "nōlīs", "nōlit", "nōlīmus", "nōlītis", "nōlint"],
|
||||
},
|
||||
"ferō": {
|
||||
("present", "ind", "active"): ["ferō", "fers", "fert", "ferimus", "fertis", "ferunt"],
|
||||
("imperfect", "ind", "active"): ["ferēbam", "ferēbās", "ferēbat", "ferēbāmus", "ferēbātis", "ferēbant"],
|
||||
("future", "ind", "active"): ["feram", "ferēs", "feret", "ferēmus", "ferētis", "ferent"],
|
||||
("perfect", "ind", "active"): ["tulī", "tulistī", "tulit", "tulimus", "tulistis", "tulērunt"],
|
||||
("present", "sbjv", "active"): ["feram", "ferās", "ferat", "ferāmus", "ferātis", "ferant"],
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def conjugate(lemma, tense, mood, voice="active", person="third", number="singular"):
|
||||
"""Return (surface, confidence). Perfect-passive forms are periphrastic and
|
||||
handled in the realizer (sum + PPP); this returns synthetic forms only."""
|
||||
lemma = lemma.strip()
|
||||
i = _idx(person, number)
|
||||
ir = _IRREG.get(lemma)
|
||||
if ir:
|
||||
tbl = ir.get((tense, mood, voice)) or ir.get((tense, mood, "active"))
|
||||
if tbl and tbl[i]:
|
||||
return tbl[i], "rule"
|
||||
v = _VERBS.get(lemma)
|
||||
if not v:
|
||||
v = _infer_principal_parts(lemma)
|
||||
if not v:
|
||||
return lemma, "fallback"
|
||||
conj, pstem, perfstem, supstem = v
|
||||
# imperative (present active) 2sg / 2pl
|
||||
if mood == "imp":
|
||||
return _imperative(conj, pstem, person, number), "rule"
|
||||
# perfect-system active
|
||||
if tense in ("perfect", "pluperfect", "futureperfect") and voice == "active":
|
||||
if not perfstem:
|
||||
return lemma, "fallback"
|
||||
end = _PERF_ACT.get((tense, mood))
|
||||
if end:
|
||||
return perfstem + end[i], "rule"
|
||||
# present-system (active + passive)
|
||||
if tense in ("present", "imperfect", "future"):
|
||||
form = _present_system(conj, pstem, tense, mood, voice, person, number)
|
||||
if form:
|
||||
return form, "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
def _imperative(conj, pstem, person, number):
|
||||
if number == "singular":
|
||||
return {1: pstem + "ā", 2: pstem + "ē", 3: pstem + "e",
|
||||
"3io": pstem + "e", 4: pstem + "ī"}[conj]
|
||||
return {1: pstem + "āte", 2: pstem + "ēte", 3: pstem + "ite",
|
||||
"3io": pstem + "ite", 4: pstem + "īte"}[conj]
|
||||
|
||||
|
||||
def _infer_principal_parts(lemma):
|
||||
"""OOV fallback: infer conjugation + stems from the 1sg-present citation form.
|
||||
Perfect/supine stems are guessed regularly (often wrong for 3rd conj) and the
|
||||
resulting forms are still returned as 'rule' but the realizer down-weights."""
|
||||
if lemma.endswith("ō"):
|
||||
base = lemma[:-1]
|
||||
# can't distinguish conj from 1sg alone reliably; default by ending vowel
|
||||
if base.endswith("i"):
|
||||
return ("3io", base[:-1], base[:-1] + "īv", base[:-1] + "īt")
|
||||
return (3, base, base + "s", base + "t")
|
||||
return None
|
||||
|
||||
|
||||
# ── PUBLIC: participles ─────────────────────────────────────────────────────────
|
||||
def participle(lemma, kind, case="nom", gender="m", number="singular"):
|
||||
"""kind: 'prs' (present active, -ns/-ntis), 'pfv' (perfect passive, -tus),
|
||||
'fut' (future active, -tūrus). Declined as an adjective via rule endings.
|
||||
Returns (form, conf)."""
|
||||
v = _VERBS.get(lemma)
|
||||
if not v:
|
||||
return lemma, "fallback"
|
||||
conj, pstem, perfstem, supstem = v
|
||||
if kind == "pfv":
|
||||
if not supstem:
|
||||
return lemma, "fallback"
|
||||
base = supstem[:-1] if supstem.endswith("t") or supstem.endswith("s") else supstem
|
||||
stem = supstem # supine stem already ends in t/s: amāt- -> amātus
|
||||
return _decline_us_a_um(stem, case, gender, number), "rule"
|
||||
if kind == "fut":
|
||||
if not supstem:
|
||||
return lemma, "fallback"
|
||||
return _decline_us_a_um(supstem + "ūr", case, gender, number), "rule"
|
||||
if kind == "prs":
|
||||
# present active participle: stem + ns (nom), stem + nt- (oblique), 3rd-decl
|
||||
pv = {1: "ā", 2: "ē", 3: "ē", "3io": "iē", 4: "iē"}[conj]
|
||||
ntstem = pstem + pv + "nt"
|
||||
return _decline_pres_ptcp(pstem + pv, case, gender, number), "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
def _decline_us_a_um(stem, case, gender, number):
|
||||
"""Decline a -us/-a/-um adjective/participle stem (2-1-2 declension)."""
|
||||
C = _CASE_MAP.get(case, case.upper())
|
||||
end = {
|
||||
("NOM", "m", "singular"): "us", ("NOM", "f", "singular"): "a", ("NOM", "n", "singular"): "um",
|
||||
("GEN", "m", "singular"): "ī", ("GEN", "f", "singular"): "ae", ("GEN", "n", "singular"): "ī",
|
||||
("DAT", "m", "singular"): "ō", ("DAT", "f", "singular"): "ae", ("DAT", "n", "singular"): "ō",
|
||||
("ACC", "m", "singular"): "um", ("ACC", "f", "singular"): "am", ("ACC", "n", "singular"): "um",
|
||||
("ABL", "m", "singular"): "ō", ("ABL", "f", "singular"): "ā", ("ABL", "n", "singular"): "ō",
|
||||
("VOC", "m", "singular"): "e", ("VOC", "f", "singular"): "a", ("VOC", "n", "singular"): "um",
|
||||
("NOM", "m", "plural"): "ī", ("NOM", "f", "plural"): "ae", ("NOM", "n", "plural"): "a",
|
||||
("GEN", "m", "plural"): "ōrum", ("GEN", "f", "plural"): "ārum", ("GEN", "n", "plural"): "ōrum",
|
||||
("DAT", "m", "plural"): "īs", ("DAT", "f", "plural"): "īs", ("DAT", "n", "plural"): "īs",
|
||||
("ACC", "m", "plural"): "ōs", ("ACC", "f", "plural"): "ās", ("ACC", "n", "plural"): "a",
|
||||
("ABL", "m", "plural"): "īs", ("ABL", "f", "plural"): "īs", ("ABL", "n", "plural"): "īs",
|
||||
("VOC", "m", "plural"): "ī", ("VOC", "f", "plural"): "ae", ("VOC", "n", "plural"): "a",
|
||||
}.get((C, gender, number), "us")
|
||||
return stem + end
|
||||
|
||||
|
||||
def _decline_pres_ptcp(stem, case, gender, number):
|
||||
"""Present active participle (amāns, amantis) — 3rd-declension, stem+ns/nt."""
|
||||
C = _CASE_MAP.get(case, case.upper())
|
||||
if C == "NOM" and number == "singular":
|
||||
return stem + "ns"
|
||||
if C == "VOC" and number == "singular":
|
||||
return stem + "ns"
|
||||
base = stem + "nt"
|
||||
end = {
|
||||
("GEN", "singular"): "is", ("DAT", "singular"): "ī",
|
||||
("ACC", "singular"): "em" if gender != "n" else "",
|
||||
("ABL", "singular"): "e",
|
||||
("NOM", "plural"): "ēs" if gender != "n" else "ia",
|
||||
("GEN", "plural"): "ium", ("DAT", "plural"): "ibus",
|
||||
("ACC", "plural"): "ēs" if gender != "n" else "ia",
|
||||
("ABL", "plural"): "ibus", ("VOC", "plural"): "ēs",
|
||||
}.get((C, number), "is")
|
||||
if C == "ACC" and number == "singular" and gender == "n":
|
||||
return stem + "ns"
|
||||
return base + end
|
||||
|
||||
|
||||
def infinitive(lemma, tense="present", voice="active"):
|
||||
lemma = lemma.strip()
|
||||
if lemma == "sum":
|
||||
return ("esse", "rule") if tense == "present" else ("fuisse", "rule")
|
||||
v = _VERBS.get(lemma)
|
||||
if not v:
|
||||
return lemma, "fallback"
|
||||
conj, pstem, perfstem, supstem = v
|
||||
if tense == "present":
|
||||
if voice == "active":
|
||||
return _active_infinitive_stem(conj, pstem).rstrip() + \
|
||||
("re" if conj != 3 and conj != "3io" else "re"), "rule"
|
||||
# passive present infinitive
|
||||
base = {1: pstem + "ā", 2: pstem + "ē", 4: pstem + "ī"}.get(conj)
|
||||
if base:
|
||||
return base + "rī", "rule"
|
||||
return pstem + "ī", "rule" # 3rd: regī
|
||||
if tense == "perfect" and voice == "active" and perfstem:
|
||||
return perfstem + "isse", "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
def lexicon_stats():
|
||||
return {
|
||||
"noun_adj_source": "UniMorph Latin (github.com/unimorph/lat, CC-BY-SA 3.0)",
|
||||
"verb_source": "rule-based 4-conjugation engine over curated attested "
|
||||
"principal parts (UniMorph verb list is a 947-lemma sample "
|
||||
"MISSING all core verbs — amō/sum/videō absent)",
|
||||
"noun_lemmas": len(_NOUNS),
|
||||
"adj_lemmas": len(_ADJS),
|
||||
"curated_verb_lemmas": len(_VERBS) + len(_IRREG),
|
||||
"gender_inference": "declension-based (nom+gen endings) + curated exceptions",
|
||||
}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
import json
|
||||
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
|
||||
print("\n-- noun declension puella (1st, fem) --")
|
||||
for c in ("nom", "gen", "dat", "acc", "abl", "voc"):
|
||||
print(f" {c}: sg={decline_noun('puella', c, 'singular')[0]:10} "
|
||||
f"pl={decline_noun('puella', c, 'plural')[0]}")
|
||||
print("\n-- rēx (3rd, m):", [decline_noun('rēx', c, 'singular')[0] for c in ('nom','gen','dat','acc','abl')])
|
||||
print("-- gender: puella=", noun_gender("puella"), "rēx=", noun_gender("rēx"),
|
||||
"bellum=", noun_gender("bellum"), "corpus=", noun_gender("corpus"),
|
||||
"manus=", noun_gender("manus"), "diēs=", noun_gender("diēs"))
|
||||
print("\n-- conjugate videō (2nd) present ind active --")
|
||||
for p in ("first", "second", "third"):
|
||||
for n in ("singular", "plural"):
|
||||
print(f" {p[:3]}.{n[:2]}: {conjugate('videō','present','ind','active',p,n)[0]}")
|
||||
print("-- amō forms:", conjugate("amō","present","ind","active","first","singular")[0],
|
||||
conjugate("amō","imperfect","ind","active","third","plural")[0],
|
||||
conjugate("amō","future","ind","active","first","singular")[0],
|
||||
conjugate("amō","perfect","ind","active","third","singular")[0])
|
||||
print("-- sum:", [conjugate("sum","present","ind","active",p,"singular")[0] for p in ("first","second","third")])
|
||||
print("-- participle amō pfv acc.f.sg:", participle("amō","pfv","acc","f","singular")[0])
|
||||
print("-- infinitive amō:", infinitive("amō")[0], "| regō pass:", infinitive("regō", voice="passive")[0])
|
||||
@@ -0,0 +1,538 @@
|
||||
"""morphology_pt_full.py — production-grade Brazilian-Portuguese morphological generator.
|
||||
|
||||
NOT a toy. Backed by two real, broad, Wiktionary-lineage lexicons:
|
||||
|
||||
VERBS — UniMorph Portuguese (github.com/unimorph/por, CC-BY-SA 3.0)
|
||||
4,001 verb lemmas × full paradigm (283,991 finite/non-finite forms +
|
||||
20,005 participle forms). Every mood/tense pt actually inflects:
|
||||
indicative present / preterite (PST;PFV) / imperfect (PST;IPFV) /
|
||||
pluperfect-simple (PST;PRF) / future,
|
||||
conditional (futuro do pretérito),
|
||||
subjunctive present / imperfect / FUTURE (PT-specific live tense),
|
||||
affirmative + negative imperative,
|
||||
PERSONAL infinitive (V;{p};{n};NFIN — a PT-specific finite-ish form),
|
||||
past participle (4 gender/number forms) + gerúndio (V.PTCP;PRS).
|
||||
|
||||
NOUNS + ADJECTIVES — kaikki.org Portuguese (Wiktionary extract, same lineage)
|
||||
81,138 noun lemmas WITH inherent gender + real (often irregular) plural —
|
||||
so -ão→-ões / -ãos / -ães / -õos is resolved PER LEMMA by Wiktionary,
|
||||
never guessed (mão→mãos, pão→pães, coração→corações).
|
||||
40,252 adjective lemmas with real feminine + masc/fem plural forms.
|
||||
|
||||
Fallbacks (degrade, never crash, on out-of-vocabulary input):
|
||||
verbs : rule generator for regular -ar/-er/-ir paradigms
|
||||
nouns : gender heuristic (endings) + rule pluralization (with -ão FLAGGED)
|
||||
adjs : -o/-a gender rule + rule pluralization
|
||||
|
||||
Confidence flag on every form:
|
||||
"lexicon" straight from UniMorph/kaikki (trust: high)
|
||||
"rule" deterministic rule (trust: medium)
|
||||
"fallback" could not inflect; returned lemma (trust: low -> FLAG)
|
||||
|
||||
Public API (used by realizer_pt.py):
|
||||
conjugate(lemma, mood, tense, person, number) -> (form, conf)
|
||||
personal_infinitive(lemma, person, number) -> (form, conf)
|
||||
participle(lemma, gender="m", number="singular") -> (form, conf)
|
||||
gerund(lemma) -> (form, conf)
|
||||
noun_gender(lemma) -> "m"|"f"
|
||||
inflect_noun(lemma, number, gender=None) -> (form, conf)
|
||||
inflect_adj(lemma, gender, number) -> (form, conf)
|
||||
lexicon_stats() -> dict
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import pickle
|
||||
|
||||
_HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
_UNIMORPH = os.path.join(_HERE, "data", "por.unimorph")
|
||||
_KAIKKI = os.path.join(_HERE, "data", "kaikki_pt.jsonl")
|
||||
_CACHE = os.path.join(_HERE, "data", "pt_morph_cache.pkl")
|
||||
|
||||
# ── mood/tense pair -> UniMorph feature triple (a in tag; b in tag; c in tag) ────
|
||||
_VERB_KEYMAP = {
|
||||
("ind", "present"): ("IND", "PRS", None),
|
||||
("ind", "preterite"): ("IND", "PST", "PFV"),
|
||||
("ind", "imperfect"): ("IND", "PST", "IPFV"),
|
||||
("ind", "pluperfect"): ("IND", "PST", "PRF"), # simple mais-que-perfeito
|
||||
("ind", "future"): ("IND", "FUT", None),
|
||||
("ind", "conditional"): ("COND", None, None),
|
||||
("sbjv", "present"): ("SBJV", "PRS", None),
|
||||
("sbjv", "imperfect"): ("SBJV", "PST", "IPFV"),
|
||||
("sbjv", "future"): ("SBJV", "FUT", None), # PT-specific
|
||||
("imp", "affirmative"): ("IMP", "POS", None),
|
||||
("imp", "negative"): ("IMP", "NEG", None),
|
||||
}
|
||||
_PERSON = {"first": "1", "second": "2", "third": "3"}
|
||||
_NUMBER = {"singular": "SG", "plural": "PL"}
|
||||
|
||||
|
||||
def _feat_set(tag):
|
||||
return set(tag.split(";"))
|
||||
|
||||
|
||||
# ── build the compact lexicon from UniMorph (verbs) + kaikki (nouns/adjs) ────────
|
||||
def _build_verbs():
|
||||
verbs = {} # (lemma, "mood|tense|person|number") -> form
|
||||
pinf = {} # (lemma, "person|number") -> personal-infinitive form
|
||||
part = {} # lemma -> {("m","SG"): form, ...} past participle
|
||||
ger = {} # lemma -> gerúndio
|
||||
with open(_UNIMORPH, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
line = line.rstrip("\n")
|
||||
if not line or "\t" not in line:
|
||||
continue
|
||||
parts = line.split("\t")
|
||||
if len(parts) != 3:
|
||||
continue
|
||||
lemma, form, tag = parts
|
||||
f = _feat_set(tag)
|
||||
head = tag.split(";")[0]
|
||||
|
||||
if head == "V.PTCP":
|
||||
if "PST" in f: # past participle: falado/falada/falados/faladas
|
||||
g = "m" if "MASC" in f else ("f" if "FEM" in f else "m")
|
||||
num = "SG" if "SG" in f else ("PL" if "PL" in f else "SG")
|
||||
part.setdefault(lemma, {})[(g, num)] = form
|
||||
elif "PRS" in f: # gerúndio: falando
|
||||
ger.setdefault(lemma, form)
|
||||
continue
|
||||
|
||||
if head != "V":
|
||||
continue
|
||||
|
||||
# personal / impersonal infinitive
|
||||
if "NFIN" in f:
|
||||
person = next((p for p in ("1", "2", "3") if p in f), None)
|
||||
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
|
||||
if person and number:
|
||||
pinf[(lemma, f"{person}|{number}")] = form
|
||||
continue
|
||||
|
||||
# finite forms
|
||||
mt = None
|
||||
for (mood, tense), (a, b, c) in _VERB_KEYMAP.items():
|
||||
if a not in f:
|
||||
continue
|
||||
if b is not None and b not in f:
|
||||
continue
|
||||
if c is not None and c not in f:
|
||||
continue
|
||||
# IND;PST needs exactly PFV|IPFV|PRF — reject if the required one absent
|
||||
mt = (mood, tense)
|
||||
break
|
||||
if mt is None:
|
||||
continue
|
||||
person = next((p for p in ("1", "2", "3") if p in f), None)
|
||||
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
|
||||
if person is None or number is None:
|
||||
continue
|
||||
verbs.setdefault((lemma, f"{mt[0]}|{mt[1]}|{person}|{number}"), form)
|
||||
return verbs, pinf, part, ger
|
||||
|
||||
|
||||
def _kaikki_gender(arg):
|
||||
if not arg:
|
||||
return None
|
||||
a = arg.lower()
|
||||
if a.startswith("f"):
|
||||
return "f"
|
||||
if a.startswith("m"):
|
||||
return "m"
|
||||
return None
|
||||
|
||||
|
||||
def _build_nouns_adjs():
|
||||
nouns = {} # lemma -> {"g","SG","PL"}
|
||||
adjs = {} # lemma -> {("m","SG"),("f","SG"),("m","PL"),("f","PL")}
|
||||
with open(_KAIKKI, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
try:
|
||||
d = json.loads(line)
|
||||
except Exception:
|
||||
continue
|
||||
pos = d.get("pos")
|
||||
word = d.get("word", "")
|
||||
if not word or " " in word: # skip multiword entries
|
||||
continue
|
||||
forms = d.get("forms", []) or []
|
||||
|
||||
if pos == "noun":
|
||||
ht = d.get("head_templates") or []
|
||||
g = None
|
||||
if ht:
|
||||
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
|
||||
if g is None:
|
||||
tags = d.get("tags") or []
|
||||
if "feminine" in tags:
|
||||
g = "f"
|
||||
elif "masculine" in tags:
|
||||
g = "m"
|
||||
pl = None
|
||||
for x in forms:
|
||||
t = x.get("tags") or []
|
||||
if "plural" in t and "alternative" not in t and "obsolete" not in t:
|
||||
pl = x.get("form")
|
||||
break
|
||||
# first entry wins; but a later entry with a plural fills a gap
|
||||
if word not in nouns:
|
||||
nouns[word] = {"g": g, "SG": word, "PL": pl}
|
||||
else:
|
||||
cur = nouns[word]
|
||||
if cur.get("g") is None and g:
|
||||
cur["g"] = g
|
||||
if not cur.get("PL") and pl:
|
||||
cur["PL"] = pl
|
||||
|
||||
elif pos == "adj":
|
||||
d0 = adjs.setdefault(word, {})
|
||||
d0.setdefault(("m", "SG"), word)
|
||||
for x in forms:
|
||||
t = set(x.get("tags") or [])
|
||||
fm = x.get("form")
|
||||
if not fm or ("alternative" in t) or ("obsolete" in t):
|
||||
continue
|
||||
if "comparative" in t or "superlative" in t or \
|
||||
"diminutive" in t or "augmentative" in t:
|
||||
continue
|
||||
if "feminine" in t and "plural" in t:
|
||||
d0[("f", "PL")] = fm
|
||||
elif "masculine" in t and "plural" in t:
|
||||
d0[("m", "PL")] = fm
|
||||
elif "feminine" in t:
|
||||
d0[("f", "SG")] = fm
|
||||
elif "plural" in t: # invariant-gender adj (feliz -> felizes)
|
||||
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
|
||||
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
|
||||
return nouns, adjs
|
||||
|
||||
|
||||
def _build_cache():
|
||||
verbs, pinf, part, ger = _build_verbs()
|
||||
nouns, adjs = _build_nouns_adjs()
|
||||
data = {"verbs": verbs, "pinf": pinf, "part": part, "ger": ger,
|
||||
"nouns": nouns, "adjs": adjs}
|
||||
try:
|
||||
with open(_CACHE, "wb") as fh:
|
||||
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
|
||||
except OSError:
|
||||
pass
|
||||
return data
|
||||
|
||||
|
||||
def _load():
|
||||
if os.path.exists(_CACHE):
|
||||
newest_src = max(os.path.getmtime(_UNIMORPH),
|
||||
os.path.getmtime(_KAIKKI) if os.path.exists(_KAIKKI) else 0)
|
||||
if os.path.getmtime(_CACHE) >= newest_src:
|
||||
try:
|
||||
with open(_CACHE, "rb") as fh:
|
||||
return pickle.load(fh)
|
||||
except Exception:
|
||||
pass
|
||||
return _build_cache()
|
||||
|
||||
|
||||
_LEX = _load()
|
||||
_VERBS, _PINF, _PART, _GER, _NOUNS, _ADJS = (
|
||||
_LEX["verbs"], _LEX["pinf"], _LEX["part"], _LEX["ger"],
|
||||
_LEX["nouns"], _LEX["adjs"])
|
||||
|
||||
|
||||
# ── regular-ending rule fallback (deterministic, last resort) ────────────────────
|
||||
def _vclass(lemma):
|
||||
return lemma[-2:] if lemma[-2:] in ("ar", "er", "ir") else None
|
||||
|
||||
|
||||
def _stem(lemma):
|
||||
return lemma[:-2]
|
||||
|
||||
|
||||
# endings indexed [1sg,2sg,3sg,1pl,2pl,3pl]
|
||||
_REG = {
|
||||
("ind", "present", "ar"): ["o", "as", "a", "amos", "ais", "am"],
|
||||
("ind", "present", "er"): ["o", "es", "e", "emos", "eis", "em"],
|
||||
("ind", "present", "ir"): ["o", "es", "e", "imos", "is", "em"],
|
||||
("ind", "preterite", "ar"): ["ei", "aste", "ou", "amos", "astes", "aram"],
|
||||
("ind", "preterite", "er"): ["i", "este", "eu", "emos", "estes", "eram"],
|
||||
("ind", "preterite", "ir"): ["i", "iste", "iu", "imos", "istes", "iram"],
|
||||
("ind", "imperfect", "ar"): ["ava", "avas", "ava", "ávamos", "áveis", "avam"],
|
||||
("ind", "imperfect", "er"): ["ia", "ias", "ia", "íamos", "íeis", "iam"],
|
||||
("ind", "imperfect", "ir"): ["ia", "ias", "ia", "íamos", "íeis", "iam"],
|
||||
("sbjv", "present", "ar"): ["e", "es", "e", "emos", "eis", "em"],
|
||||
("sbjv", "present", "er"): ["a", "as", "a", "amos", "ais", "am"],
|
||||
("sbjv", "present", "ir"): ["a", "as", "a", "amos", "ais", "am"],
|
||||
("sbjv", "imperfect", "ar"): ["asse", "asses", "asse", "ássemos", "ásseis", "assem"],
|
||||
("sbjv", "imperfect", "er"): ["esse", "esses", "esse", "êssemos", "êsseis", "essem"],
|
||||
("sbjv", "imperfect", "ir"): ["isse", "isses", "isse", "íssemos", "ísseis", "issem"],
|
||||
("sbjv", "future", "ar"): ["ar", "ares", "ar", "armos", "ardes", "arem"],
|
||||
("sbjv", "future", "er"): ["er", "eres", "er", "ermos", "erdes", "erem"],
|
||||
("sbjv", "future", "ir"): ["ir", "ires", "ir", "irmos", "irdes", "irem"],
|
||||
}
|
||||
# future & conditional attach to the FULL infinitive
|
||||
_FUT = ["ei", "ás", "á", "emos", "eis", "ão"]
|
||||
_COND = ["ia", "ias", "ia", "íamos", "íeis", "iam"]
|
||||
|
||||
|
||||
def _slot_idx(person, number):
|
||||
base = {"first": 0, "second": 1, "third": 2}[person]
|
||||
return base + (0 if number == "singular" else 3)
|
||||
|
||||
|
||||
def _rule_conjugate(lemma, mood, tense, person, number):
|
||||
vc = _vclass(lemma)
|
||||
if vc is None:
|
||||
return None
|
||||
st, i = _stem(lemma), _slot_idx(person, number)
|
||||
if mood == "ind" and tense == "future":
|
||||
return lemma + _FUT[i]
|
||||
if mood == "ind" and tense == "conditional":
|
||||
return lemma + _COND[i]
|
||||
if mood == "imp": # affirmative tú/vocês imperative ~ subjunctive present
|
||||
table = _REG.get(("sbjv", "present", vc))
|
||||
if table and tense == "negative":
|
||||
return st + table[i]
|
||||
# affirmative 2sg = 3sg present indicative; others = subjunctive
|
||||
pres = _REG.get(("ind", "present", vc))
|
||||
if person == "second" and number == "singular":
|
||||
return st + pres[2]
|
||||
return st + table[i] if table else None
|
||||
table = _REG.get((mood, tense, vc))
|
||||
if table:
|
||||
return st + table[i]
|
||||
return None
|
||||
|
||||
|
||||
# verified corrections to UniMorph data errors (each audited individually, not
|
||||
# guessed). The three 1PL-present entries are glued-allomorph errors surfaced by a
|
||||
# full-lexicon scan for a non-final "mos" in V;1;PL;IND;PRS forms (the ONLY three).
|
||||
_VERB_FIX = {
|
||||
("estar", "ind", "imperfect", "third", "plural"): "estavam", # was "estávam"
|
||||
("estar", "ind", "present", "first", "plural"): "estamos", # was "estamosestámos"
|
||||
("haver", "ind", "present", "first", "plural"): "havemos", # was "havemoshemos"
|
||||
("ir", "ind", "present", "first", "plural"): "vamos", # was "vamosimos"
|
||||
}
|
||||
|
||||
|
||||
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
|
||||
def conjugate(lemma, mood, tense, person, number):
|
||||
"""Return (surface, confidence). mood in ind|sbjv|imp; tense per _VERB_KEYMAP."""
|
||||
lemma = lemma.strip().lower()
|
||||
fix = _VERB_FIX.get((lemma, mood, tense, person, number))
|
||||
if fix:
|
||||
return fix, "lexicon"
|
||||
p, n = _PERSON.get(person), _NUMBER.get(number)
|
||||
if p and n:
|
||||
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
|
||||
if form:
|
||||
# pt-BR normalization: UniMorph `por` carries the EUROPEAN spelling of
|
||||
# the -ar 1pl PRETERITE (-ámos). Brazilian PT drops the accent
|
||||
# (falámos->falamos, chegámos->chegamos) — 3,334/4,001 verbs affected.
|
||||
if (mood == "ind" and tense == "preterite" and person == "first"
|
||||
and number == "plural" and form.endswith("ámos")):
|
||||
form = form[:-4] + "amos"
|
||||
return form, "lexicon"
|
||||
r = _rule_conjugate(lemma, mood, tense, person, number)
|
||||
if r:
|
||||
return r, "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
def personal_infinitive(lemma, person, number):
|
||||
"""PT personal (inflected) infinitive: para falarmos, ao chegarem."""
|
||||
lemma = lemma.strip().lower()
|
||||
p, n = _PERSON.get(person), _NUMBER.get(number)
|
||||
if p and n:
|
||||
form = _PINF.get((lemma, f"{p}|{n}"))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
# rule: infinitive + personal endings (-, -es, -, -mos, -des, -em)
|
||||
end = {("first", "singular"): "", ("second", "singular"): "es",
|
||||
("third", "singular"): "", ("first", "plural"): "mos",
|
||||
("second", "plural"): "des", ("third", "plural"): "em"}.get((person, number), "")
|
||||
return lemma + end, "rule"
|
||||
|
||||
|
||||
# ── PUBLIC: participle + gerund ───────────────────────────────────────────────────
|
||||
def participle(lemma, gender="m", number="singular"):
|
||||
lemma = lemma.strip().lower()
|
||||
g = "f" if gender == "f" else "m"
|
||||
num = "SG" if number == "singular" else "PL"
|
||||
d = _PART.get(lemma)
|
||||
if d:
|
||||
form = d.get((g, num)) or d.get(("m", "SG"))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
if lemma.endswith("ar"):
|
||||
base = lemma[:-2] + "ad"
|
||||
elif lemma[-2:] in ("er", "ir"):
|
||||
base = lemma[:-2] + "id"
|
||||
else:
|
||||
return lemma, "fallback"
|
||||
suf = {"m|SG": "o", "f|SG": "a", "m|PL": "os", "f|PL": "as"}[f"{g}|{num}"]
|
||||
return base + suf, "rule"
|
||||
|
||||
|
||||
def gerund(lemma):
|
||||
lemma = lemma.strip().lower()
|
||||
if lemma in _GER:
|
||||
return _GER[lemma], "lexicon"
|
||||
if lemma.endswith("ar"):
|
||||
return lemma[:-2] + "ando", "rule"
|
||||
if lemma.endswith("er"):
|
||||
return lemma[:-2] + "endo", "rule"
|
||||
if lemma.endswith("ir"):
|
||||
return lemma[:-2] + "indo", "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
|
||||
_FEM_SUF = ("ção", "são", "ção", "dade", "tade", "agem", "igem", "ugem", "gem",
|
||||
"ez", "eza", "ice", "ície", "tude", "ude", "âncbefore")
|
||||
_FEM_SUF = ("ção", "são", "dade", "tade", "agem", "gem", "eza", "ez", "ice",
|
||||
"tude", "ude", "ância", "ência", "ínia")
|
||||
_MASC_SUF = ("ema", "oma", "ama", "grama", "eta", "ão") # Greek -ma etc. (mostly m)
|
||||
|
||||
|
||||
def _gender_heuristic(noun):
|
||||
for suf in _FEM_SUF:
|
||||
if noun.endswith(suf):
|
||||
return "f"
|
||||
if noun.endswith(("ema", "oma", "ama")): # problema, idioma, programa
|
||||
return "m"
|
||||
if noun.endswith("a") or noun.endswith("ã"):
|
||||
return "f"
|
||||
if noun.endswith("o") or noun.endswith(("l", "r", "z", "m", "u", "i")):
|
||||
return "m"
|
||||
return "m"
|
||||
|
||||
|
||||
def noun_gender(lemma):
|
||||
lemma = lemma.strip().lower()
|
||||
d = _NOUNS.get(lemma)
|
||||
if d and d.get("g"):
|
||||
return d["g"]
|
||||
return _gender_heuristic(lemma)
|
||||
|
||||
|
||||
_INVARIANT_PL_SUF = ("s",) # paroxytones ending -s are invariant (o lápis / os lápis)
|
||||
|
||||
|
||||
def _rule_plural(noun):
|
||||
"""Deterministic PT pluralization. Returns (form, ok) where ok=False flags an
|
||||
ambiguous -ão that should lower confidence (the lexicon normally resolves it)."""
|
||||
if not noun:
|
||||
return noun, True
|
||||
if noun.endswith("ão"):
|
||||
return noun[:-2] + "ões", False # majority rule, but AMBIGUOUS -> flag
|
||||
if noun.endswith("m"):
|
||||
return noun[:-1] + "ns", True # homem->homens, jardim->jardins
|
||||
if noun.endswith("al"):
|
||||
return noun[:-2] + "ais", True
|
||||
if noun.endswith("el"):
|
||||
return noun[:-2] + "éis", True
|
||||
if noun.endswith("ol"):
|
||||
return noun[:-2] + "óis", True
|
||||
if noun.endswith("ul"):
|
||||
return noun[:-2] + "uis", True
|
||||
if noun.endswith("il"):
|
||||
return noun[:-2] + "is", True # stressed (funil->funis); unstressed rarer
|
||||
if noun.endswith(("r", "z")):
|
||||
return noun + "es", True # flor->flores, luz->luzes
|
||||
if noun.endswith("s"):
|
||||
# paroxytone -s (lápis, ônibus) invariant; oxytone -s (país) -> -es
|
||||
return noun, True
|
||||
if noun.endswith(("a", "e", "i", "o", "u", "á", "é", "í", "ó", "ú", "ã")):
|
||||
return noun + "s", True
|
||||
return noun + "s", True
|
||||
|
||||
|
||||
def inflect_noun(lemma, number, gender=None):
|
||||
lemma = lemma.strip().lower()
|
||||
d = _NOUNS.get(lemma)
|
||||
if number == "singular":
|
||||
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
|
||||
if d and d.get("PL"):
|
||||
return d["PL"], "lexicon"
|
||||
form, ok = _rule_plural(lemma)
|
||||
return form, ("rule" if ok else "fallback")
|
||||
|
||||
|
||||
# ── PUBLIC: adjective agreement ──────────────────────────────────────────────────
|
||||
def inflect_adj(lemma, gender, number):
|
||||
lemma = lemma.strip().lower()
|
||||
g = "f" if gender == "f" else "m"
|
||||
num = "SG" if number == "singular" else "PL"
|
||||
d = _ADJS.get(lemma)
|
||||
if d:
|
||||
form = d.get((g, num))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
# build a missing plural from this gender's singular
|
||||
sg = d.get((g, "SG")) or d.get(("m", "SG")) or lemma
|
||||
if num == "PL":
|
||||
pl, ok = _rule_plural(sg)
|
||||
return pl, ("rule" if ok else "fallback")
|
||||
return sg, "lexicon"
|
||||
# rule fallback: -o/-a gender, then pluralize
|
||||
a = lemma
|
||||
if g == "f":
|
||||
if a.endswith("o"):
|
||||
a = a[:-1] + "a"
|
||||
elif a.endswith(("ês", "or")) and not a.endswith("ior"):
|
||||
a = a + "a" # português->portuguesa, trabalhador->..a
|
||||
if num == "PL":
|
||||
a, ok = _rule_plural(a)
|
||||
return a, ("rule" if ok else "fallback")
|
||||
return a, "rule"
|
||||
|
||||
|
||||
def lexicon_stats():
|
||||
return {
|
||||
"verb_source": "UniMorph Portuguese (github.com/unimorph/por)",
|
||||
"noun_adj_source": "kaikki.org Portuguese (Wiktionary extract)",
|
||||
"license": "CC-BY-SA (Wiktionary-derived)",
|
||||
"verb_forms": len(_VERBS),
|
||||
"verb_lemmas": len({k[0] for k in _VERBS}),
|
||||
"personal_infinitive_forms": len(_PINF),
|
||||
"participle_lemmas": len(_PART),
|
||||
"gerund_lemmas": len(_GER),
|
||||
"noun_lemmas": len(_NOUNS),
|
||||
"adj_lemmas": len(_ADJS),
|
||||
}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
|
||||
tests = [
|
||||
("falar", "ind", "present", "first", "singular", "falo"),
|
||||
("comer", "ind", "present", "third", "plural", "comem"),
|
||||
("partir", "ind", "present", "first", "plural", "partimos"),
|
||||
("ser", "ind", "present", "third", "singular", "é"),
|
||||
("ir", "ind", "preterite", "first", "singular", "fui"),
|
||||
("ter", "ind", "future", "first", "singular", "terei"),
|
||||
("fazer", "sbjv", "present", "first", "singular", "faça"),
|
||||
("dormir", "ind", "present", "first", "singular", "durmo"),
|
||||
("dar", "ind", "preterite", "third", "singular", "deu"),
|
||||
("poder", "ind", "conditional", "first", "singular", "poderia"),
|
||||
("fazer", "sbjv", "future", "third", "singular", "fizer"),
|
||||
("estar", "ind", "present", "third", "singular", "está"),
|
||||
]
|
||||
ok = 0
|
||||
for lemma, mood, tense, per, num, exp in tests:
|
||||
got, conf = conjugate(lemma, mood, tense, per, num)
|
||||
flag = "OK " if got == exp else "XX "
|
||||
ok += got == exp
|
||||
print(f" {flag}{lemma:8} {mood}/{tense} {per[:3]}.{num[:2]} -> {got:14} ({conf}) exp={exp}")
|
||||
print(f"verb tests {ok}/{len(tests)}")
|
||||
print(" gender: casa=", noun_gender("casa"), "problema=", noun_gender("problema"),
|
||||
"mão=", noun_gender("mão"), "coração=", noun_gender("coração"),
|
||||
"flor=", noun_gender("flor"))
|
||||
print(" plural: mão->", inflect_noun("mão", "plural"),
|
||||
"| pão->", inflect_noun("pão", "plural"),
|
||||
"| animal->", inflect_noun("animal", "plural"),
|
||||
"| coração->", inflect_noun("coração", "plural"))
|
||||
print(" adj: bonito/f/sg->", inflect_adj("bonito", "f", "singular"),
|
||||
"| feliz/m/pl->", inflect_adj("feliz", "m", "plural"),
|
||||
"| português/f/sg->", inflect_adj("português", "f", "singular"))
|
||||
print(" part: fazer/m/sg->", participle("fazer"), "| ger falar->", gerund("falar"))
|
||||
print(" pinf falar 1pl->", personal_infinitive("falar", "first", "plural"))
|
||||
@@ -0,0 +1,609 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""morphology_ro_full.py — production-grade Romanian morphological generator.
|
||||
|
||||
Romanian is the BIG typological delta of the Romance family. The verb engine and
|
||||
the confidence/fallback contract TRANSFER from the Italian sibling; the NOMINAL
|
||||
system is genuinely new: Romanian has a SUFFIXED definite article, a preserved
|
||||
NOM/ACC vs GEN/DAT case distinction, a NEUTER gender (masc-agreeing in SG,
|
||||
fem-agreeing in PL), and a VOCATIVE. Those are grounded in real per-lemma data,
|
||||
not guessed.
|
||||
|
||||
Real, Wiktionary-lineage lexical sources:
|
||||
|
||||
VERBS — UniMorph Romanian (github.com/unimorph/ron, CC-BY-SA 3.0)
|
||||
~1216 verb lemmas × paradigm, CLEAN orthography:
|
||||
indicativ prezent / imperfect (PST;IPFV) / perfectul simplu (PST;PFV) /
|
||||
conjunctiv prezent (SBJV;PRS, stored WITHOUT the 'să' particle),
|
||||
participiu (V.PTCP;PST, INVARIABLE in the perfect compus),
|
||||
gerunziu (V.CVB;PRS), infinitiv (NFIN), imperativ.
|
||||
ro_irreg_verbs (embedded) — high-frequency verbs UniMorph MISSES
|
||||
(avea, vrea, da) + the auxiliary clitic paradigms the compound tenses need
|
||||
(perfect-compus am/ai/a/am/ați/au, viitor voi/vei/va/vom/veți/vor,
|
||||
condițional aș/ai/ar/am/ați/ar). Real standard forms.
|
||||
|
||||
NOUNS — kaikki.org Romanian (Wiktionary extract, CC-BY-SA 3.0)
|
||||
the FULL declension per lemma, cleanly tagged:
|
||||
(nom/acc | gen/dat | vocative) × (indefinite | definite) × (sg | pl).
|
||||
This is what makes the suffixed article LEXICALLY grounded (om→omul,
|
||||
casă→casa, băiat→băiatul, casei gen/dat, omule vocative). Inherent gender
|
||||
m / f / n (NEUTER available directly) from the head template.
|
||||
|
||||
ADJECTIVES — UniMorph Romanian ADJ
|
||||
full case × gender(MASC/FEM/NEUT) × number × definiteness paradigm.
|
||||
|
||||
Fallbacks (degrade, never crash, on OOV): rule verb conjugation for -a/-ea/-e/-i/-î
|
||||
classes, rule pluralization, rule suffixed-article by gender+ending. Every form
|
||||
carries a confidence flag: "lexicon" | "rule" | "fallback".
|
||||
|
||||
Public API (used by realizer_ro.py):
|
||||
conjugate(lemma, mood, tense, person, number) -> (form, conf)
|
||||
aux(kind, person, number) -> str # perfect / future / conditional clitics
|
||||
participle(lemma) -> (form, conf) # INVARIABLE
|
||||
gerund(lemma) -> (form, conf)
|
||||
noun_gender(lemma) -> "m"|"f"|"n"
|
||||
definite_suffix(noun, gender, number, case) -> (form, conf) # rule engine
|
||||
inflect_noun(lemma, number, gender=None, case="nomacc", definite=False) -> (form, conf)
|
||||
inflect_adj(lemma, gender, number, case="nomacc", definite=False) -> (form, conf)
|
||||
lexicon_stats() -> dict
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import pickle
|
||||
|
||||
_HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
_UNIMORPH = os.path.join(_HERE, "data", "ron.unimorph")
|
||||
_KAIKKI = os.path.join(_HERE, "data", "kaikki_ro.jsonl")
|
||||
_CACHE = os.path.join(_HERE, "data", "ro_morph_cache.pkl")
|
||||
|
||||
# ── (mood, tense) -> UniMorph feature set ─────────────────────────────────────────
|
||||
_VERB_KEYMAP = {
|
||||
("ind", "present"): {"IND", "PRS"},
|
||||
("ind", "imperfect"): {"IND", "PST", "IPFV"},
|
||||
("ind", "perfect_s"): {"IND", "PST", "PFV"}, # perfectul simplu (regional/lit.)
|
||||
("sbjv", "present"): {"SBJV", "PRS"},
|
||||
("imp", "affirmative"): {"POS", "IMP"},
|
||||
}
|
||||
_PERSON = {"first": "1", "second": "2", "third": "3"}
|
||||
_NUMBER = {"singular": "SG", "plural": "PL"}
|
||||
|
||||
|
||||
def _feat_set(tag):
|
||||
return set(tag.split(";"))
|
||||
|
||||
|
||||
# ── high-frequency irregulars UniMorph misses + auxiliary clitic paradigms ────────
|
||||
# Real standard Romanian forms (textbook paradigms).
|
||||
_IRREG = {
|
||||
"avea": {
|
||||
"ind|present|1|SG": "am", "ind|present|2|SG": "ai", "ind|present|3|SG": "are",
|
||||
"ind|present|1|PL": "avem", "ind|present|2|PL": "aveți", "ind|present|3|PL": "au",
|
||||
"ind|imperfect|1|SG": "aveam", "ind|imperfect|2|SG": "aveai",
|
||||
"ind|imperfect|3|SG": "avea", "ind|imperfect|1|PL": "aveam",
|
||||
"ind|imperfect|2|PL": "aveați", "ind|imperfect|3|PL": "aveau",
|
||||
"sbjv|present|3|SG": "aibă", "sbjv|present|3|PL": "aibă",
|
||||
"sbjv|present|1|SG": "am", "sbjv|present|2|SG": "ai",
|
||||
"sbjv|present|1|PL": "avem", "sbjv|present|2|PL": "aveți",
|
||||
"part": "avut", "ger": "având",
|
||||
},
|
||||
"vrea": {
|
||||
"ind|present|1|SG": "vreau", "ind|present|2|SG": "vrei", "ind|present|3|SG": "vrea",
|
||||
"ind|present|1|PL": "vrem", "ind|present|2|PL": "vreți", "ind|present|3|PL": "vor",
|
||||
"ind|imperfect|1|SG": "voiam", "ind|imperfect|3|SG": "voia",
|
||||
"sbjv|present|3|SG": "vrea", "sbjv|present|3|PL": "vrea",
|
||||
"part": "vrut", "ger": "vrând",
|
||||
},
|
||||
"da": {
|
||||
"ind|present|1|SG": "dau", "ind|present|2|SG": "dai", "ind|present|3|SG": "dă",
|
||||
"ind|present|1|PL": "dăm", "ind|present|2|PL": "dați", "ind|present|3|PL": "dau",
|
||||
"ind|imperfect|1|SG": "dădeam", "ind|imperfect|3|SG": "dădea",
|
||||
"sbjv|present|3|SG": "dea", "sbjv|present|3|PL": "dea",
|
||||
"part": "dat", "ger": "dând",
|
||||
},
|
||||
"fi": { # a fi — present is in UniMorph but keep participle + subjunctive here
|
||||
"part": "fost", "ger": "fiind",
|
||||
"sbjv|present|1|SG": "fiu", "sbjv|present|2|SG": "fii", "sbjv|present|3|SG": "fie",
|
||||
"sbjv|present|1|PL": "fim", "sbjv|present|2|PL": "fiți", "sbjv|present|3|PL": "fie",
|
||||
"ind|imperfect|1|SG": "eram", "ind|imperfect|2|SG": "erai",
|
||||
"ind|imperfect|3|SG": "era", "ind|imperfect|1|PL": "eram",
|
||||
"ind|imperfect|2|PL": "erați", "ind|imperfect|3|PL": "erau",
|
||||
},
|
||||
}
|
||||
# auxiliary clitic paradigms (person,number)->form
|
||||
_AUX = {
|
||||
"perfect": {("first", "singular"): "am", ("second", "singular"): "ai",
|
||||
("third", "singular"): "a", ("first", "plural"): "am",
|
||||
("second", "plural"): "ați", ("third", "plural"): "au"},
|
||||
"future": {("first", "singular"): "voi", ("second", "singular"): "vei",
|
||||
("third", "singular"): "va", ("first", "plural"): "vom",
|
||||
("second", "plural"): "veți", ("third", "plural"): "vor"},
|
||||
"conditional": {("first", "singular"): "aș", ("second", "singular"): "ai",
|
||||
("third", "singular"): "ar", ("first", "plural"): "am",
|
||||
("second", "plural"): "ați", ("third", "plural"): "ar"},
|
||||
}
|
||||
|
||||
|
||||
def aux(kind, person, number):
|
||||
return _AUX[kind][(person, number)]
|
||||
|
||||
|
||||
# ── build verb lexicon from UniMorph ──────────────────────────────────────────────
|
||||
def _build_verbs():
|
||||
verbs, part, ger = {}, {}, {}
|
||||
with open(_UNIMORPH, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
line = line.rstrip("\n")
|
||||
if not line or "\t" not in line:
|
||||
continue
|
||||
parts = line.split("\t")
|
||||
if len(parts) != 3:
|
||||
continue
|
||||
lemma, form, tag = parts
|
||||
f = _feat_set(tag)
|
||||
head = tag.split(";")[0]
|
||||
if head == "V.PTCP":
|
||||
if "PST" in f:
|
||||
part.setdefault(lemma, form)
|
||||
continue
|
||||
if head == "V.CVB":
|
||||
if "PRS" in f:
|
||||
ger.setdefault(lemma, form)
|
||||
continue
|
||||
if head != "V":
|
||||
continue
|
||||
person = next((p for p in ("1", "2", "3") if p in f), None)
|
||||
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
|
||||
if person is None or number is None:
|
||||
continue
|
||||
# conjunctiv forms in UniMorph carry a leading 'să ' — strip it
|
||||
surf = form
|
||||
if surf.startswith("să "):
|
||||
surf = surf[3:]
|
||||
for (mood, tense), req in _VERB_KEYMAP.items():
|
||||
if not req <= f:
|
||||
continue
|
||||
if tense == "imperfect" and "PFV" in f:
|
||||
continue
|
||||
if tense == "perfect_s" and "IPFV" in f:
|
||||
continue
|
||||
# keep IND;PRS out of the PRF slot (mai-mult-ca-perfect etc. ignored)
|
||||
if {"IND", "PRS"} <= req and "PRF" in f:
|
||||
continue
|
||||
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), surf)
|
||||
break
|
||||
return verbs, part, ger
|
||||
|
||||
|
||||
# ── kaikki nouns: full declension paradigm per lemma ──────────────────────────────
|
||||
_EXCL = {"alternative", "archaic", "obsolete", "regional", "dialectal", "rare",
|
||||
"table-tags", "inflection-template", "error-unrecognized-form",
|
||||
"diminutive", "augmentative", "informal"}
|
||||
|
||||
|
||||
def _noun_key(tagset):
|
||||
if tagset & _EXCL:
|
||||
return None
|
||||
if "vocative" in tagset:
|
||||
case = "voc"
|
||||
elif "genitive" in tagset or "dative" in tagset:
|
||||
case = "gendat"
|
||||
elif "nominative" in tagset or "accusative" in tagset:
|
||||
case = "nomacc"
|
||||
else:
|
||||
return None
|
||||
definite = "definite" in tagset and "indefinite" not in tagset
|
||||
number = "PL" if "plural" in tagset else ("SG" if "singular" in tagset else None)
|
||||
if number is None:
|
||||
return None
|
||||
return (case, definite, number)
|
||||
|
||||
|
||||
def _build_nouns():
|
||||
nouns = {} # lemma -> {"g":..., para:{(case,def,num):form}, "PL":plain_plural}
|
||||
with open(_KAIKKI, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
try:
|
||||
d = json.loads(line)
|
||||
except Exception:
|
||||
continue
|
||||
if d.get("pos") != "noun":
|
||||
continue
|
||||
word = d.get("word", "")
|
||||
if not word or " " in word:
|
||||
continue
|
||||
ht = d.get("head_templates") or []
|
||||
g = None
|
||||
if ht:
|
||||
a = str((ht[0].get("args") or {}).get("1") or "").lower()
|
||||
if a[:1] in ("m", "f", "n"):
|
||||
g = a[:1]
|
||||
entry = nouns.setdefault(word, {"g": g, "para": {}, "PL": None})
|
||||
if entry["g"] is None and g:
|
||||
entry["g"] = g
|
||||
for x in (d.get("forms") or []):
|
||||
fm = x.get("form")
|
||||
tg = set(x.get("tags") or [])
|
||||
if not fm or fm in ("-", "#", "") or " " in fm:
|
||||
continue
|
||||
if tg == {"plural"} and not entry["PL"]:
|
||||
entry["PL"] = fm
|
||||
k = _noun_key(tg)
|
||||
if k and k not in entry["para"]:
|
||||
entry["para"][k] = fm
|
||||
return nouns
|
||||
|
||||
|
||||
# ── adjectives from kaikki (UniMorph ron ADJ is sparse AND mis-tagged; kaikki is
|
||||
# clean: the 4-form agreement pattern bun/bună/buni/bune). Neuter maps sg->masc,
|
||||
# pl->fem, so 4 forms (m/f × SG/PL) fully cover it. ────────────────────────────
|
||||
def _build_adjs():
|
||||
adjs = {} # lemma -> {(gender,number): form} gender in {m,f}
|
||||
with open(_KAIKKI, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
try:
|
||||
d = json.loads(line)
|
||||
except Exception:
|
||||
continue
|
||||
if d.get("pos") != "adj":
|
||||
continue
|
||||
word = d.get("word", "")
|
||||
if not word or " " in word:
|
||||
continue
|
||||
d0 = adjs.setdefault(word, {})
|
||||
d0.setdefault(("m", "SG"), word) # masc sg = headword
|
||||
for x in (d.get("forms") or []):
|
||||
fm = x.get("form")
|
||||
t = set(x.get("tags") or [])
|
||||
if not fm or " " in fm or fm in ("-", "#") or (t & _EXCL):
|
||||
continue
|
||||
if "definite" in t or "genitive" in t or "dative" in t:
|
||||
continue # keep indefinite nom/acc agr set
|
||||
pl = "plural" in t
|
||||
fem = "feminine" in t
|
||||
masc = "masculine" in t
|
||||
if fem and pl:
|
||||
d0.setdefault(("f", "PL"), fm)
|
||||
elif masc and pl:
|
||||
d0.setdefault(("m", "PL"), fm)
|
||||
elif fem and not pl:
|
||||
d0.setdefault(("f", "SG"), fm)
|
||||
elif pl and not fem and not masc: # bare plural -> both genders
|
||||
d0.setdefault(("m", "PL"), fm)
|
||||
d0.setdefault(("f", "PL"), fm)
|
||||
return adjs
|
||||
|
||||
|
||||
def _build_cache():
|
||||
verbs, part, ger = _build_verbs()
|
||||
nouns = _build_nouns()
|
||||
adjs = _build_adjs()
|
||||
data = {"verbs": verbs, "part": part, "ger": ger, "nouns": nouns, "adjs": adjs}
|
||||
try:
|
||||
with open(_CACHE, "wb") as fh:
|
||||
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
|
||||
except OSError:
|
||||
pass
|
||||
return data
|
||||
|
||||
|
||||
def _load():
|
||||
if os.path.exists(_CACHE):
|
||||
srcs = [_UNIMORPH, _KAIKKI]
|
||||
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
|
||||
if os.path.getmtime(_CACHE) >= newest:
|
||||
try:
|
||||
with open(_CACHE, "rb") as fh:
|
||||
return pickle.load(fh)
|
||||
except Exception:
|
||||
pass
|
||||
return _build_cache()
|
||||
|
||||
|
||||
_LEX = _load()
|
||||
_VERBS, _PART, _GER, _NOUNS, _ADJS = (
|
||||
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"])
|
||||
|
||||
|
||||
# ── rule verb conjugation fallback ────────────────────────────────────────────────
|
||||
def _vclass(lemma):
|
||||
if lemma.endswith("a"):
|
||||
return "a"
|
||||
if lemma.endswith("ea"):
|
||||
return "ea"
|
||||
if lemma.endswith("e"):
|
||||
return "e"
|
||||
if lemma.endswith("i"):
|
||||
return "i"
|
||||
if lemma.endswith("î"):
|
||||
return "î"
|
||||
return None
|
||||
|
||||
|
||||
# regular present endings by class [1sg,2sg,3sg,1pl,2pl,3pl]
|
||||
_REG_PRS = {
|
||||
"a": ["", "i", "ă", "ăm", "ați", "ă"], # a lucra type (simplified)
|
||||
"ea": ["", "i", "e", "em", "eți", "", ],
|
||||
"e": ["", "i", "e", "em", "eți", ""],
|
||||
"i": ["esc", "ești", "ește", "im", "iți", "esc"], # -i type (a vorbi)
|
||||
"î": ["ăsc", "ăști", "ăște", "âm", "âți", "ăsc"],
|
||||
}
|
||||
_SLOT = {("first", "singular"): 0, ("second", "singular"): 1, ("third", "singular"): 2,
|
||||
("first", "plural"): 3, ("second", "plural"): 4, ("third", "plural"): 5}
|
||||
|
||||
|
||||
def _rule_conjugate(lemma, mood, tense, person, number):
|
||||
vc = _vclass(lemma)
|
||||
if vc is None:
|
||||
return None
|
||||
i = _SLOT[(person, number)]
|
||||
body = lemma[:-len(vc)]
|
||||
if mood == "ind" and tense == "present":
|
||||
end = _REG_PRS[vc][i]
|
||||
return body + end
|
||||
if mood == "ind" and tense == "imperfect":
|
||||
# -a/-i/-î -> stem + a/eai...; -e/-ea -> eam. Simplified regular imperfect.
|
||||
stem = body
|
||||
endings = {"a": ["am", "ai", "a", "am", "ați", "au"],
|
||||
"i": ["eam", "eai", "ea", "eam", "eați", "eau"],
|
||||
"î": ["am", "ai", "a", "am", "ați", "au"],
|
||||
"e": ["eam", "eai", "ea", "eam", "eați", "eau"],
|
||||
"ea": ["eam", "eai", "ea", "eam", "eați", "eau"]}[vc]
|
||||
return stem + endings[i]
|
||||
return None
|
||||
|
||||
|
||||
# ── PUBLIC verb API ───────────────────────────────────────────────────────────────
|
||||
def conjugate(lemma, mood, tense, person, number):
|
||||
lemma = lemma.strip().lower()
|
||||
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{_NUMBER.get(number,'?')}"
|
||||
ir = _IRREG.get(lemma)
|
||||
if ir and key in ir:
|
||||
return ir[key], "lexicon"
|
||||
form = _VERBS.get((lemma, key))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
r = _rule_conjugate(lemma, mood, tense, person, number)
|
||||
if r is not None:
|
||||
return r, "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
def participle(lemma):
|
||||
"""Past participle — INVARIABLE in the perfect compus (am mers, am văzut)."""
|
||||
lemma = lemma.strip().lower()
|
||||
ir = _IRREG.get(lemma)
|
||||
if ir and "part" in ir:
|
||||
return ir["part"], "lexicon"
|
||||
if lemma in _PART:
|
||||
return _PART[lemma], "lexicon"
|
||||
vc = _vclass(lemma)
|
||||
if vc == "a":
|
||||
return lemma[:-1] + "at", "rule"
|
||||
if vc in ("ea",):
|
||||
return lemma[:-2] + "ut", "rule"
|
||||
if vc == "i":
|
||||
return lemma[:-1] + "it", "rule"
|
||||
if vc == "î":
|
||||
return lemma[:-1] + "ât", "rule"
|
||||
if vc == "e":
|
||||
return lemma[:-1] + "ut", "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
def gerund(lemma):
|
||||
lemma = lemma.strip().lower()
|
||||
ir = _IRREG.get(lemma)
|
||||
if ir and "ger" in ir:
|
||||
return ir["ger"], "lexicon"
|
||||
if lemma in _GER:
|
||||
return _GER[lemma], "lexicon"
|
||||
vc = _vclass(lemma)
|
||||
if vc in ("a", "î"):
|
||||
return lemma[:-1] + "ând", "rule"
|
||||
if vc in ("ea", "e", "i"):
|
||||
return lemma[:-len(vc)] + "ind", "rule"
|
||||
return lemma, "fallback"
|
||||
|
||||
|
||||
# ── noun gender ───────────────────────────────────────────────────────────────────
|
||||
def noun_gender(lemma):
|
||||
lemma = lemma.strip().lower()
|
||||
d = _NOUNS.get(lemma)
|
||||
if d and d.get("g") in ("m", "f", "n"):
|
||||
return d["g"]
|
||||
if lemma.endswith(("ă", "a", "e")):
|
||||
return "f"
|
||||
return "m"
|
||||
|
||||
|
||||
# ── SUFFIXED DEFINITE ARTICLE — rule engine (fallback for OOV nouns) ───────────────
|
||||
def definite_suffix(noun, gender, number, case="nomacc"):
|
||||
"""Attach the enclitic definite article by gender + ending. Returns (form, conf).
|
||||
This is the headline Romanian-specific engine extension."""
|
||||
n = noun
|
||||
g = gender
|
||||
if number == "singular":
|
||||
if g in ("m", "n"):
|
||||
if case == "gendat":
|
||||
# masc/neut gen-dat definite: -lui
|
||||
if n.endswith("e"):
|
||||
return n + "lui", "rule" # câine -> câinelui
|
||||
if n.endswith("u"):
|
||||
return n + "lui", "rule"
|
||||
return n + "ului", "rule" # om -> omului
|
||||
# nom/acc
|
||||
if n.endswith("e"):
|
||||
return n + "le", "rule" # câine -> câinele
|
||||
if n.endswith("u"):
|
||||
return n + "l", "rule" # codru -> codrul
|
||||
if n.endswith("i"):
|
||||
return n + "ul", "rule"
|
||||
return n + "ul", "rule" # om -> omul
|
||||
# feminine singular
|
||||
if case == "gendat":
|
||||
# fem gen/dat definite = plural-stem + i (casei, fetei) — needs plural;
|
||||
# approximated as: -ă->-ei, -e->-ei, -a->-alei
|
||||
if n.endswith("ă"):
|
||||
return n[:-1] + "ei", "rule" # casă -> casei
|
||||
if n.endswith("e"):
|
||||
return n[:-1] + "ei", "rule" # carte -> cărții(approx cartei)
|
||||
if n.endswith("a"):
|
||||
return n[:-1] + "lei", "rule"
|
||||
return n + "i", "rule"
|
||||
# fem nom/acc
|
||||
if n.endswith("ă"):
|
||||
return n[:-1] + "a", "rule" # casă -> casa
|
||||
if n.endswith("e"):
|
||||
return n[:-1] + "ea", "rule" # carte -> cartea
|
||||
if n.endswith("a"):
|
||||
return n + "ua", "rule" # stea -> steaua
|
||||
if n.endswith("i"):
|
||||
return n + "a", "rule"
|
||||
return n + "a", "rule"
|
||||
# plural
|
||||
if case == "gendat":
|
||||
base = noun
|
||||
return base + "lor", "rule" # -lor for all gen/dat pl
|
||||
if g == "m":
|
||||
return noun + "i", "rule" # oameni -> oamenii (+i)
|
||||
return noun + "le", "rule" # case -> casele, trenuri->trenurile
|
||||
|
||||
|
||||
# ── rule pluralization (fallback) ─────────────────────────────────────────────────
|
||||
def _rule_plural(noun, gender):
|
||||
if gender == "f":
|
||||
if noun.endswith("ă"):
|
||||
return noun[:-1] + "e"
|
||||
if noun.endswith("e"):
|
||||
return noun[:-1] + "i"
|
||||
if noun.endswith("a"):
|
||||
return noun[:-1] + "le"
|
||||
return noun + "e"
|
||||
if gender == "n":
|
||||
return noun + "uri"
|
||||
# masculine
|
||||
if noun.endswith(("e",)):
|
||||
return noun[:-1] + "i"
|
||||
return noun + "i"
|
||||
|
||||
|
||||
# ── PUBLIC noun inflection ────────────────────────────────────────────────────────
|
||||
def inflect_noun(lemma, number, gender=None, case="nomacc", definite=False):
|
||||
lemma = lemma.strip().lower()
|
||||
g = gender or noun_gender(lemma)
|
||||
d = _NOUNS.get(lemma)
|
||||
numk = "SG" if number == "singular" else "PL"
|
||||
if d:
|
||||
if case == "voc":
|
||||
form = d["para"].get(("voc", True, numk)) or d["para"].get(("voc", False, numk))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
# try the exact paradigm cell from kaikki (lexically grounded)
|
||||
form = d["para"].get((case, definite, numk))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
# indefinite fallbacks from the paradigm
|
||||
if not definite:
|
||||
form = d["para"].get(("nomacc", False, numk))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
if numk == "PL" and d.get("PL"):
|
||||
return d["PL"], "lexicon"
|
||||
if numk == "SG":
|
||||
return lemma, "lexicon"
|
||||
# rule path
|
||||
base = lemma if number == "singular" else _rule_plural(lemma, g)
|
||||
if definite:
|
||||
return definite_suffix(base, g, number, case)
|
||||
return base, ("rule" if d is None else "lexicon")
|
||||
|
||||
|
||||
# ── PUBLIC adjective agreement ────────────────────────────────────────────────────
|
||||
def _neuter_map(gender, number):
|
||||
# neuter agrees masculine in SG, feminine in PL
|
||||
if gender == "n":
|
||||
return "m" if number == "singular" else "f"
|
||||
return gender
|
||||
|
||||
|
||||
def inflect_adj(lemma, gender, number, case="nomacc", definite=False):
|
||||
lemma = lemma.strip().lower()
|
||||
numk = "SG" if number == "singular" else "PL"
|
||||
eg = _neuter_map(gender, number) # neuter -> masc(SG)/fem(PL)
|
||||
d = _ADJS.get(lemma)
|
||||
if d:
|
||||
form = d.get((eg, numk))
|
||||
if form:
|
||||
return form, "lexicon"
|
||||
# rule fallback: 4-form pattern bun/bună/buni/bune keyed by effective gender
|
||||
a = lemma
|
||||
if number == "singular":
|
||||
if eg == "f":
|
||||
if a.endswith("e"):
|
||||
return a, "rule" # mare invariant sg
|
||||
if a.endswith("u"):
|
||||
return a[:-1] + "ă", "rule" # nou -> nouă
|
||||
if a.endswith("ă"):
|
||||
return a, "rule"
|
||||
return a + "ă", "rule" # bun -> bună
|
||||
return a, "rule" # masc/neut sg = lemma
|
||||
# plural
|
||||
if eg == "f":
|
||||
if a.endswith("e"):
|
||||
return a[:-1] + "i", "rule" # mare -> mari
|
||||
if a.endswith("u"):
|
||||
return a[:-1] + "e", "rule" # nou -> noue (approx; 'noi' irr)
|
||||
if a.endswith("ă"):
|
||||
return a[:-1] + "e", "rule"
|
||||
return a + "e", "rule" # bun -> bune
|
||||
# masc/neut(SG-only)->here masc pl -> -i
|
||||
if a.endswith("e"):
|
||||
return a[:-1] + "i", "rule" # mare -> mari
|
||||
if a.endswith("u"):
|
||||
return a[:-1] + "i", "rule"
|
||||
return a + "i", "rule" # bun -> buni
|
||||
|
||||
|
||||
def lexicon_stats():
|
||||
return {
|
||||
"verb_source": "UniMorph Romanian (github.com/unimorph/ron) + curated "
|
||||
"irregulars (avea/vrea/da + aux clitic paradigms)",
|
||||
"noun_source": "kaikki.org Romanian — full case/definite/vocative declension",
|
||||
"adj_source": "UniMorph Romanian ADJ (case×gender×number×definiteness)",
|
||||
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
|
||||
"unimorph_verb_forms": len(_VERBS),
|
||||
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
|
||||
"irregular_verb_lemmas": len(_IRREG),
|
||||
"participle_lemmas": len(_PART),
|
||||
"noun_lemmas": len(_NOUNS),
|
||||
"adj_lemmas": len(_ADJS),
|
||||
}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
|
||||
print("\n── SUFFIXED DEFINITE ARTICLE (the headline delta) ──")
|
||||
for n, g in [("om", "m"), ("băiat", "m"), ("casă", "f"), ("carte", "f"),
|
||||
("tren", "n"), ("student", "m"), ("floare", "f")]:
|
||||
sg = inflect_noun(n, "singular", g, "nomacc", True)
|
||||
pl = inflect_noun(n, "plural", g, "nomacc", True)
|
||||
gd = inflect_noun(n, "singular", g, "gendat", True)
|
||||
vo = inflect_noun(n, "singular", g, "voc", False)
|
||||
print(f" {n:8}({g}) def.sg={sg[0]:12} def.pl={pl[0]:14} "
|
||||
f"gen/dat.sg={gd[0]:12} voc={vo[0]}")
|
||||
print("\n── NEUTER split agreement (tren: masc SG / fem PL) ──")
|
||||
print(" tren nou ->", inflect_noun("tren", "singular", "n")[0],
|
||||
inflect_adj("nou", "n", "singular")[0])
|
||||
print(" trenuri noi->", inflect_noun("tren", "plural", "n")[0],
|
||||
inflect_adj("nou", "n", "plural")[0])
|
||||
print("\n── verbs ──")
|
||||
for l, m, t, p, n, in [("merge", "ind", "present", "third", "singular"),
|
||||
("avea", "ind", "present", "first", "singular"),
|
||||
("fi", "ind", "present", "third", "singular"),
|
||||
("vorbi", "ind", "present", "third", "plural"),
|
||||
("face", "sbjv", "present", "third", "singular"),
|
||||
("lucra", "ind", "imperfect", "third", "singular")]:
|
||||
print(f" {l:8}{m}/{t:10}{p[:3]}.{n[:2]} -> {conjugate(l,m,t,p,n)}")
|
||||
print(" perfect-aux(3sg):", aux("perfect", "third", "singular"),
|
||||
"| future(1sg):", aux("future", "first", "singular"),
|
||||
"| cond(3sg):", aux("conditional", "third", "singular"))
|
||||
print(" participle merge/vedea:", participle("merge"), participle("vedea"))
|
||||
@@ -0,0 +1,43 @@
|
||||
// multilingual_gate.el - deterministic language detect + localized-phrase test.
|
||||
|
||||
fn mg_det(text: String, want: String) -> String {
|
||||
let got: String = ml_detect(text)
|
||||
let ok: String = "MISMATCH"
|
||||
if str_eq(got, want) { let ok = "ok" }
|
||||
return " detect(" + got + ") want=" + want + " (" + ok + ") :: " + text + "\n"
|
||||
}
|
||||
|
||||
fn mg_ok(text: String, want: String) -> Int {
|
||||
if str_eq(ml_detect(text), want) { return 1 }
|
||||
return 0
|
||||
}
|
||||
|
||||
fn run_ml_gate() -> String {
|
||||
let t1: String = "Does Neuron use SQLite for storage?"
|
||||
let t2: String = "Neuron, me explica cómo la saliencia forma las geometrías."
|
||||
let t3: String = "O professor não leu o livro na memória."
|
||||
let t4: String = "Che cosa memorizza Neuron nella memoria?"
|
||||
|
||||
let rep: String = "==== ELP multilingual detect + localized phrases ====\n"
|
||||
let rep = rep + mg_det(t1, "en")
|
||||
let rep = rep + mg_det(t2, "es")
|
||||
let rep = rep + mg_det(t3, "pt")
|
||||
let rep = rep + mg_det(t4, "it")
|
||||
|
||||
let rep = rep + " localized decline (pt): " + ml_tr("no_memory", "pt") + "\n"
|
||||
let rep = rep + " localized decline (es): " + ml_tr("no_memory", "es") + "\n"
|
||||
let rep = rep + " term(saliência->en): " + ml_term("saliência", "pt") + "\n"
|
||||
let rep = rep + " pred(store->pt): " + ml_translate_pred("store", "pt") + "\n"
|
||||
|
||||
let ok: Int = 0
|
||||
if mg_ok(t1, "en") == 1 { let ok = ok + 1 }
|
||||
if mg_ok(t2, "es") == 1 { let ok = ok + 1 }
|
||||
if mg_ok(t3, "pt") == 1 { let ok = ok + 1 }
|
||||
if mg_ok(t4, "it") == 1 { let ok = ok + 1 }
|
||||
let rep = rep + "-----------------------------------------------------------------\n"
|
||||
let rep = rep + "language detected correctly: " + int_to_str(ok) + "/4\n"
|
||||
if ok == 4 { let rep = rep + "ML GATE: PASS\n" } else { let rep = rep + "ML GATE: FAIL\n" }
|
||||
return rep
|
||||
}
|
||||
|
||||
println(run_ml_gate())
|
||||
@@ -0,0 +1,52 @@
|
||||
// propositions_gate.el - the READ primitive over memory text (native el).
|
||||
// Proves triples are recovered from free memory text and that SACRED polarity
|
||||
// survives extraction (a negative memory must yield a NOT-triple).
|
||||
|
||||
fn pg_check(text: String, want_pol: String) -> String {
|
||||
let p: [String] = prop_extract_one(text, "nd-test")
|
||||
let pol: String = slots_get(p, "polarity")
|
||||
let ok: String = "MISMATCH"
|
||||
if str_eq(pol, want_pol) { let ok = "ok" }
|
||||
return " " + prop_repr(p) + " pol=" + pol + " expected=" + want_pol + " (" + ok + ")\n"
|
||||
}
|
||||
|
||||
fn pg_pol_ok(text: String, want_pol: String) -> Int {
|
||||
let p: [String] = prop_extract_one(text, "nd-test")
|
||||
if str_eq(slots_get(p, "polarity"), want_pol) { return 1 }
|
||||
return 0
|
||||
}
|
||||
|
||||
fn run_prop_gate() -> String {
|
||||
let m1: String = "Neuron stores memories in SQLite."
|
||||
let m2: String = "The engram does not delete a memory."
|
||||
let m3: String = "Salience never drops the negation."
|
||||
let m4: String = "The teacher gives the book to the children."
|
||||
|
||||
let rep: String = "==== ELP proposition extraction (memory text -> triples) ====\n"
|
||||
let rep = rep + pg_check(m1, "aff")
|
||||
let rep = rep + pg_check(m2, "neg")
|
||||
let rep = rep + pg_check(m3, "neg")
|
||||
let rep = rep + pg_check(m4, "aff")
|
||||
|
||||
// multi-sentence memory: one triple per sentence, order preserved
|
||||
let doc: String = "Neuron persists learning. It does not forget the library."
|
||||
let props: [String] = prop_extract(doc, "nd-doc")
|
||||
let rep = rep + " --- multi-sentence doc (" + int_to_str(native_list_len(props)) + " props) ---\n"
|
||||
let di: Int = 0
|
||||
while di < native_list_len(props) {
|
||||
let rep = rep + " " + native_list_get(props, di) + "\n"
|
||||
let di = di + 1
|
||||
}
|
||||
|
||||
let ok: Int = 0
|
||||
if pg_pol_ok(m1, "aff") == 1 { let ok = ok + 1 }
|
||||
if pg_pol_ok(m2, "neg") == 1 { let ok = ok + 1 }
|
||||
if pg_pol_ok(m3, "neg") == 1 { let ok = ok + 1 }
|
||||
if pg_pol_ok(m4, "aff") == 1 { let ok = ok + 1 }
|
||||
let rep = rep + "-----------------------------------------------------------------\n"
|
||||
let rep = rep + "SACRED polarity correct on extraction: " + int_to_str(ok) + "/4\n"
|
||||
if ok == 4 { let rep = rep + "PROP GATE: PASS\n" } else { let rep = rep + "PROP GATE: FAIL\n" }
|
||||
return rep
|
||||
}
|
||||
|
||||
println(run_prop_gate())
|
||||
Vendored
BIN
Binary file not shown.
Vendored
+254
-105
@@ -10,6 +10,9 @@ el_val_t query_param(el_val_t path, el_val_t key);
|
||||
el_val_t query_int(el_val_t path, el_val_t key, el_val_t default_val);
|
||||
el_val_t extract_id(el_val_t path, el_val_t prefix);
|
||||
el_val_t route_stats(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_act_stats(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_text_health(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t persist_canonical(void);
|
||||
el_val_t route_create_node(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_get_node(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_scan_nodes(el_val_t method, el_val_t path, el_val_t body);
|
||||
@@ -17,21 +20,29 @@ el_val_t route_scan_edges(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_search(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_activate(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_create_edge(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_create_edges_batch(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_neighbors(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_strengthen(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_forget(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_create_ise(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_sync(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_save(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_load(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_health(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_embed_backfill(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_sync(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_load_merge(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_emit_ise(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_capture_knowledge(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t route_similarity(el_val_t method, el_val_t path, el_val_t body);
|
||||
el_val_t check_auth_ok(el_val_t method, el_val_t body);
|
||||
el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body);
|
||||
|
||||
el_val_t bind_raw;
|
||||
el_val_t bind_str;
|
||||
el_val_t port;
|
||||
el_val_t data_dir_raw;
|
||||
el_val_t data_dir;
|
||||
el_val_t snapshot_path;
|
||||
el_val_t boot_snap;
|
||||
|
||||
el_val_t parse_port(el_val_t bind) {
|
||||
el_val_t colon = str_index_of(bind, EL_STR(":"));
|
||||
@@ -110,17 +121,40 @@ el_val_t route_stats(el_val_t method, el_val_t path, el_val_t body) {
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_act_stats(el_val_t method, el_val_t path, el_val_t body) {
|
||||
return engram_act_stats_json();
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_text_health(el_val_t method, el_val_t path, el_val_t body) {
|
||||
return engram_text_health_json();
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t persist_canonical(void) {
|
||||
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
|
||||
el_val_t dir = ({ el_val_t _if_result_1 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_1 = (EL_STR("/tmp/engram")); } else { _if_result_1 = (dir_raw); } _if_result_1; });
|
||||
return engram_save(el_str_concat(dir, EL_STR("/snapshot.json")));
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_create_node(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t content = json_get_string(body, EL_STR("content"));
|
||||
el_val_t node_type = json_get_string(body, EL_STR("node_type"));
|
||||
if (str_eq(node_type, EL_STR(""))) {
|
||||
node_type = EL_STR("Memory");
|
||||
}
|
||||
el_val_t salience = json_get_float(body, EL_STR("salience"));
|
||||
if (salience == el_from_float(0.0)) {
|
||||
salience = el_from_float(0.5);
|
||||
}
|
||||
el_val_t id = engram_node(content, node_type, salience);
|
||||
el_val_t nt_raw = json_get_string(body, EL_STR("node_type"));
|
||||
el_val_t node_type = ({ el_val_t _if_result_2 = 0; if (str_eq(nt_raw, EL_STR(""))) { _if_result_2 = (EL_STR("Memory")); } else { _if_result_2 = (nt_raw); } _if_result_2; });
|
||||
el_val_t sal_present = json_get_raw(body, EL_STR("salience"));
|
||||
el_val_t salience = ({ el_val_t _if_result_3 = 0; if (str_eq(sal_present, EL_STR(""))) { _if_result_3 = (el_from_float(0.5)); } else { _if_result_3 = (json_get_float(body, EL_STR("salience"))); } _if_result_3; });
|
||||
el_val_t label_raw = json_get_string(body, EL_STR("label"));
|
||||
el_val_t label = ({ el_val_t _if_result_4 = 0; if (str_eq(label_raw, EL_STR(""))) { _if_result_4 = (content); } else { _if_result_4 = (label_raw); } _if_result_4; });
|
||||
el_val_t imp_present = json_get_raw(body, EL_STR("importance"));
|
||||
el_val_t importance = ({ el_val_t _if_result_5 = 0; if (str_eq(imp_present, EL_STR(""))) { _if_result_5 = (el_from_float(0.5)); } else { _if_result_5 = (json_get_float(body, EL_STR("importance"))); } _if_result_5; });
|
||||
el_val_t conf_present = json_get_raw(body, EL_STR("confidence"));
|
||||
el_val_t confidence = ({ el_val_t _if_result_6 = 0; if (str_eq(conf_present, EL_STR(""))) { _if_result_6 = (el_from_float(1.0)); } else { _if_result_6 = (json_get_float(body, EL_STR("confidence"))); } _if_result_6; });
|
||||
el_val_t tier_raw = json_get_string(body, EL_STR("tier"));
|
||||
el_val_t tier = ({ el_val_t _if_result_7 = 0; if (str_eq(tier_raw, EL_STR(""))) { _if_result_7 = (EL_STR("Working")); } else { _if_result_7 = (tier_raw); } _if_result_7; });
|
||||
el_val_t tags = json_get_string(body, EL_STR("tags"));
|
||||
el_val_t id = engram_node_full(content, node_type, label, salience, importance, confidence, tier, tags);
|
||||
el_val_t saved = persist_canonical();
|
||||
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"id\":\""), id), EL_STR("\",\"content\":\"")), content), EL_STR("\",\"node_type\":\"")), node_type), EL_STR("\"}"));
|
||||
return 0;
|
||||
}
|
||||
@@ -146,11 +180,9 @@ el_val_t route_scan_nodes(el_val_t method, el_val_t path, el_val_t body) {
|
||||
}
|
||||
|
||||
el_val_t route_scan_edges(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t dir = env(EL_STR("ENGRAM_DATA_DIR"));
|
||||
if (str_eq(dir, EL_STR(""))) {
|
||||
dir = EL_STR("/tmp/engram");
|
||||
}
|
||||
el_val_t snap_path = el_str_concat(dir, EL_STR("/snapshot.json"));
|
||||
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
|
||||
el_val_t dir = ({ el_val_t _if_result_8 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_8 = (EL_STR("/tmp/engram")); } else { _if_result_8 = (dir_raw); } _if_result_8; });
|
||||
el_val_t snap_path = el_str_concat(dir, EL_STR("/.scan-export.json"));
|
||||
engram_save(snap_path);
|
||||
el_val_t snap = fs_read(snap_path);
|
||||
if (str_eq(snap, EL_STR(""))) {
|
||||
@@ -165,36 +197,22 @@ el_val_t route_scan_edges(el_val_t method, el_val_t path, el_val_t body) {
|
||||
}
|
||||
|
||||
el_val_t route_search(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t q = EL_STR("");
|
||||
if (str_eq(method, EL_STR("GET"))) {
|
||||
q = query_param(path, EL_STR("q"));
|
||||
} else {
|
||||
q = json_get_string(body, EL_STR("query"));
|
||||
}
|
||||
el_val_t limit = query_int(path, EL_STR("limit"), 20);
|
||||
if (limit == 0) {
|
||||
limit = json_get_int(body, EL_STR("limit"));
|
||||
}
|
||||
if (limit == 0) {
|
||||
limit = 20;
|
||||
}
|
||||
el_val_t q = ({ el_val_t _if_result_9 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_9 = (query_param(path, EL_STR("q"))); } else { _if_result_9 = (json_get_string(body, EL_STR("query"))); } _if_result_9; });
|
||||
el_val_t lim_url = query_int(path, EL_STR("limit"), 0);
|
||||
el_val_t lim_body = json_get_int(body, EL_STR("limit"));
|
||||
el_val_t lim_either = ({ el_val_t _if_result_10 = 0; if ((lim_url > 0)) { _if_result_10 = (lim_url); } else { _if_result_10 = (lim_body); } _if_result_10; });
|
||||
el_val_t limit = ({ el_val_t _if_result_11 = 0; if ((lim_either > 0)) { _if_result_11 = (lim_either); } else { _if_result_11 = (20); } _if_result_11; });
|
||||
return engram_search_json(q, limit);
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_activate(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t q = EL_STR("");
|
||||
el_val_t depth = 3;
|
||||
if (str_eq(method, EL_STR("GET"))) {
|
||||
q = query_param(path, EL_STR("q"));
|
||||
depth = query_int(path, EL_STR("depth"), 3);
|
||||
} else {
|
||||
q = json_get_string(body, EL_STR("query"));
|
||||
el_val_t bd = json_get_int(body, EL_STR("depth"));
|
||||
if (bd > 0) {
|
||||
depth = bd;
|
||||
}
|
||||
el_val_t q = ({ el_val_t _if_result_12 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_12 = (query_param(path, EL_STR("q"))); } else { _if_result_12 = (json_get_string(body, EL_STR("query"))); } _if_result_12; });
|
||||
if (str_eq(q, EL_STR(""))) {
|
||||
return err_json(EL_STR("missing query"));
|
||||
}
|
||||
el_val_t d_raw = ({ el_val_t _if_result_13 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_13 = (query_int(path, EL_STR("depth"), 3)); } else { _if_result_13 = (json_get_int(body, EL_STR("depth"))); } _if_result_13; });
|
||||
el_val_t depth = ({ el_val_t _if_result_14 = 0; if ((d_raw > 0)) { _if_result_14 = (d_raw); } else { _if_result_14 = (3); } _if_result_14; });
|
||||
return el_str_concat(el_str_concat(EL_STR("{\"results\":"), engram_activate_json(q, depth)), EL_STR("}"));
|
||||
return 0;
|
||||
}
|
||||
@@ -202,19 +220,51 @@ el_val_t route_activate(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t route_create_edge(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t from_id = json_get_string(body, EL_STR("from_id"));
|
||||
el_val_t to_id = json_get_string(body, EL_STR("to_id"));
|
||||
el_val_t relation = json_get_string(body, EL_STR("relation"));
|
||||
if (str_eq(relation, EL_STR(""))) {
|
||||
relation = EL_STR("associates");
|
||||
}
|
||||
el_val_t weight = json_get_float(body, EL_STR("weight"));
|
||||
if (weight == el_from_float(0.0)) {
|
||||
weight = el_from_float(0.5);
|
||||
}
|
||||
el_val_t rel_raw = json_get_string(body, EL_STR("relation"));
|
||||
el_val_t relation = ({ el_val_t _if_result_15 = 0; if (str_eq(rel_raw, EL_STR(""))) { _if_result_15 = (EL_STR("associates")); } else { _if_result_15 = (rel_raw); } _if_result_15; });
|
||||
el_val_t w_present = json_get_raw(body, EL_STR("weight"));
|
||||
el_val_t weight = ({ el_val_t _if_result_16 = 0; if (str_eq(w_present, EL_STR(""))) { _if_result_16 = (el_from_float(0.5)); } else { _if_result_16 = (json_get_float(body, EL_STR("weight"))); } _if_result_16; });
|
||||
engram_connect(from_id, to_id, weight, relation);
|
||||
el_val_t saved = persist_canonical();
|
||||
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"from_id\":\""), from_id), EL_STR("\",\"to_id\":\"")), to_id), EL_STR("\",\"relation\":\"")), relation), EL_STR("\"}"));
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_create_edges_batch(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t arr = json_get_raw(body, EL_STR("edges"));
|
||||
if (str_eq(arr, EL_STR(""))) {
|
||||
return err_json(EL_STR("missing edges array"));
|
||||
}
|
||||
el_val_t n = json_array_len(arr);
|
||||
if (n == 0) {
|
||||
return EL_STR("{\"ok\":true,\"accepted\":0,\"skipped\":0}");
|
||||
}
|
||||
el_val_t i = 0;
|
||||
el_val_t accepted = 0;
|
||||
el_val_t skipped = 0;
|
||||
while (i < n) {
|
||||
el_val_t item = json_array_get(arr, i);
|
||||
el_val_t from_id = json_get_string(item, EL_STR("from_id"));
|
||||
el_val_t to_id = json_get_string(item, EL_STR("to_id"));
|
||||
if (str_eq(from_id, EL_STR("")) || str_eq(to_id, EL_STR(""))) {
|
||||
skipped = (skipped + 1);
|
||||
} else {
|
||||
el_val_t rel_raw = json_get_string(item, EL_STR("relation"));
|
||||
el_val_t relation = ({ el_val_t _if_result_17 = 0; if (str_eq(rel_raw, EL_STR(""))) { _if_result_17 = (EL_STR("associates")); } else { _if_result_17 = (rel_raw); } _if_result_17; });
|
||||
el_val_t w_present = json_get_raw(item, EL_STR("weight"));
|
||||
el_val_t weight = ({ el_val_t _if_result_18 = 0; if (str_eq(w_present, EL_STR(""))) { _if_result_18 = (el_from_float(0.5)); } else { _if_result_18 = (json_get_float(item, EL_STR("weight"))); } _if_result_18; });
|
||||
engram_connect(from_id, to_id, weight, relation);
|
||||
accepted = (accepted + 1);
|
||||
}
|
||||
i = (i + 1);
|
||||
}
|
||||
if (accepted > 0) {
|
||||
el_val_t saved = persist_canonical();
|
||||
}
|
||||
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"accepted\":"), int_to_str(accepted)), EL_STR(",\"skipped\":")), int_to_str(skipped)), EL_STR("}"));
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_neighbors(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t id = extract_id(path, EL_STR("/api/neighbors/"));
|
||||
if (str_eq(id, EL_STR(""))) {
|
||||
@@ -231,6 +281,7 @@ el_val_t route_strengthen(el_val_t method, el_val_t path, el_val_t body) {
|
||||
return err_json(EL_STR("missing node_id"));
|
||||
}
|
||||
engram_strengthen(id);
|
||||
el_val_t saved = persist_canonical();
|
||||
return ok_json();
|
||||
return 0;
|
||||
}
|
||||
@@ -241,11 +292,83 @@ el_val_t route_forget(el_val_t method, el_val_t path, el_val_t body) {
|
||||
return err_json(EL_STR("missing id"));
|
||||
}
|
||||
engram_forget(id);
|
||||
el_val_t saved = persist_canonical();
|
||||
return ok_json();
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_create_ise(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t route_save(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t p_raw = json_get_string(body, EL_STR("path"));
|
||||
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
|
||||
el_val_t dir = ({ el_val_t _if_result_19 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_19 = (EL_STR("/tmp/engram")); } else { _if_result_19 = (dir_raw); } _if_result_19; });
|
||||
el_val_t p = ({ el_val_t _if_result_20 = 0; if (str_eq(p_raw, EL_STR(""))) { _if_result_20 = (el_str_concat(dir, EL_STR("/snapshot.json"))); } else { _if_result_20 = (p_raw); } _if_result_20; });
|
||||
el_val_t sv = engram_save(p);
|
||||
el_val_t sv_ok = ({ el_val_t _if_result_21 = 0; if ((sv == 0)) { _if_result_21 = (EL_STR("false")); } else { _if_result_21 = (EL_STR("true")); } _if_result_21; });
|
||||
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":"), sv_ok), EL_STR(",\"path\":\"")), p), EL_STR("\",\"node_count\":")), int_to_str(engram_node_count())), EL_STR(",\"edge_count\":")), int_to_str(engram_edge_count())), EL_STR("}"));
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_load(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t p_raw = json_get_string(body, EL_STR("path"));
|
||||
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
|
||||
el_val_t dir = ({ el_val_t _if_result_22 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_22 = (EL_STR("/tmp/engram")); } else { _if_result_22 = (dir_raw); } _if_result_22; });
|
||||
el_val_t p = ({ el_val_t _if_result_23 = 0; if (str_eq(p_raw, EL_STR(""))) { _if_result_23 = (el_str_concat(dir, EL_STR("/snapshot.json"))); } else { _if_result_23 = (p_raw); } _if_result_23; });
|
||||
el_val_t ld = engram_load(p);
|
||||
el_val_t ld_ok = ({ el_val_t _if_result_24 = 0; if ((ld == 0)) { _if_result_24 = (EL_STR("false")); } else { _if_result_24 = (EL_STR("true")); } _if_result_24; });
|
||||
el_val_t nc_after = engram_node_count();
|
||||
el_val_t hollow = ({ el_val_t _if_result_25 = 0; if ((nc_after == 0)) { _if_result_25 = (EL_STR("true")); } else { _if_result_25 = (EL_STR("false")); } _if_result_25; });
|
||||
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":"), ld_ok), EL_STR(",\"path\":\"")), p), EL_STR("\",\"node_count\":")), int_to_str(nc_after)), EL_STR(",\"edge_count\":")), int_to_str(engram_edge_count())), EL_STR(",\"hollow\":")), hollow), EL_STR("}"));
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_health(el_val_t method, el_val_t path, el_val_t body) {
|
||||
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"status\":\"ok\",\"engine\":\"engram-runtime-native\",\"node_count\":"), int_to_str(engram_node_count())), EL_STR(",\"edge_count\":")), int_to_str(engram_edge_count())), EL_STR("}"));
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_embed_backfill(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t n = query_int(path, EL_STR("n"), 32);
|
||||
el_val_t result = engram_embed_backfill(n);
|
||||
el_val_t done = json_get_float(result, EL_STR("embedded"));
|
||||
if (done > el_from_float(0.0)) {
|
||||
el_val_t saved = persist_canonical();
|
||||
}
|
||||
return result;
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_sync(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
|
||||
el_val_t dir = ({ el_val_t _if_result_26 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_26 = (EL_STR("/tmp/engram")); } else { _if_result_26 = (dir_raw); } _if_result_26; });
|
||||
el_val_t snap_path = el_str_concat(dir, EL_STR("/.sync-export.json"));
|
||||
engram_save(snap_path);
|
||||
el_val_t snap = fs_read(snap_path);
|
||||
if (str_eq(snap, EL_STR(""))) {
|
||||
return err_json(EL_STR("sync export failed: snapshot unreadable"));
|
||||
}
|
||||
return snap;
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_load_merge(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t p = json_get_string(body, EL_STR("path"));
|
||||
if (str_eq(p, EL_STR(""))) {
|
||||
return err_json(EL_STR("path is required"));
|
||||
}
|
||||
if (str_eq(fs_read(p), EL_STR(""))) {
|
||||
return err_json(EL_STR("file missing or empty"));
|
||||
}
|
||||
el_val_t before_n = engram_node_count();
|
||||
el_val_t before_e = engram_edge_count();
|
||||
engram_load_merge(p);
|
||||
el_val_t added_n = (engram_node_count() - before_n);
|
||||
el_val_t added_e = (engram_edge_count() - before_e);
|
||||
el_val_t saved = persist_canonical();
|
||||
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"nodes_added\":"), int_to_str(added_n)), EL_STR(",\"edges_added\":")), int_to_str(added_e)), EL_STR(",\"node_count\":")), int_to_str(engram_node_count())), EL_STR("}"));
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_emit_ise(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t content = json_get_string(body, EL_STR("content"));
|
||||
if (str_eq(content, EL_STR(""))) {
|
||||
return err_json(EL_STR("missing content"));
|
||||
@@ -254,55 +377,55 @@ el_val_t route_create_ise(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t imp = el_from_float(0.3);
|
||||
el_val_t conf = el_from_float(0.8);
|
||||
el_val_t id = engram_node_full(content, EL_STR("InternalStateEvent"), EL_STR("state-event"), sal, imp, conf, EL_STR("Episodic"), EL_STR("[\"internal-state\",\"InternalStateEvent\"]"));
|
||||
el_val_t ret_raw = env(EL_STR("ENGRAM_ISE_RETENTION_MS"));
|
||||
el_val_t ret_ms = ({ el_val_t _if_result_27 = 0; if (str_eq(ret_raw, EL_STR(""))) { _if_result_27 = (172800000); } else { _if_result_27 = (str_to_int(ret_raw)); } _if_result_27; });
|
||||
el_val_t pruned = engram_prune_telemetry(ret_ms);
|
||||
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"id\":\""), id), EL_STR("\",\"pruned\":")), int_to_str(pruned)), EL_STR("}"));
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_capture_knowledge(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t content = json_get_string(body, EL_STR("content"));
|
||||
if (str_eq(content, EL_STR(""))) {
|
||||
return err_json(EL_STR("missing content"));
|
||||
}
|
||||
el_val_t title = json_get_string(body, EL_STR("title"));
|
||||
el_val_t label = ({ el_val_t _if_result_28 = 0; if (str_eq(title, EL_STR(""))) { _if_result_28 = (str_slice(content, 0, 60)); } else { _if_result_28 = (title); } _if_result_28; });
|
||||
el_val_t category_raw = json_get_string(body, EL_STR("category"));
|
||||
el_val_t category = ({ el_val_t _if_result_29 = 0; if (str_eq(category_raw, EL_STR(""))) { _if_result_29 = (EL_STR("other")); } else { _if_result_29 = (category_raw); } _if_result_29; });
|
||||
el_val_t ktier_raw = json_get_string(body, EL_STR("tier"));
|
||||
el_val_t ktier = ({ el_val_t _if_result_30 = 0; if (str_eq(ktier_raw, EL_STR(""))) { _if_result_30 = (EL_STR("note")); } else { _if_result_30 = (ktier_raw); } _if_result_30; });
|
||||
el_val_t project = json_get_string(body, EL_STR("project"));
|
||||
el_val_t tags_raw = json_get_raw(body, EL_STR("tags"));
|
||||
el_val_t tags_base = ({ el_val_t _if_result_31 = 0; if (str_eq(tags_raw, EL_STR(""))) { _if_result_31 = (EL_STR("[]")); } else { _if_result_31 = (tags_raw); } _if_result_31; });
|
||||
el_val_t base_len = str_len(tags_base);
|
||||
el_val_t head = str_slice(tags_base, 0, (base_len - 1));
|
||||
el_val_t sep = ({ el_val_t _if_result_32 = 0; if (str_eq(head, EL_STR("["))) { _if_result_32 = (EL_STR("")); } else { _if_result_32 = (EL_STR(",")); } _if_result_32; });
|
||||
el_val_t safe_cat = str_replace(category, EL_STR("\""), EL_STR("'"));
|
||||
el_val_t safe_tier = str_replace(ktier, EL_STR("\""), EL_STR("'"));
|
||||
el_val_t safe_proj = str_replace(project, EL_STR("\""), EL_STR("'"));
|
||||
el_val_t proj_tag = ({ el_val_t _if_result_33 = 0; if (str_eq(safe_proj, EL_STR(""))) { _if_result_33 = (EL_STR("")); } else { _if_result_33 = (el_str_concat(el_str_concat(EL_STR(",\"project:"), safe_proj), EL_STR("\""))); } _if_result_33; });
|
||||
el_val_t tags = el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(head, sep), EL_STR("\"category:")), safe_cat), EL_STR("\",\"tier:")), safe_tier), EL_STR("\"")), proj_tag), EL_STR("]"));
|
||||
el_val_t sal = el_from_float(0.5);
|
||||
el_val_t imp = el_from_float(0.5);
|
||||
el_val_t conf = el_from_float(0.9);
|
||||
el_val_t id = engram_node_full(content, EL_STR("Knowledge"), label, sal, imp, conf, EL_STR("Semantic"), tags);
|
||||
el_val_t saved = persist_canonical();
|
||||
return el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"id\":\""), id), EL_STR("\"}"));
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_sync(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t dir = env(EL_STR("ENGRAM_DATA_DIR"));
|
||||
if (str_eq(dir, EL_STR(""))) {
|
||||
dir = EL_STR("/tmp/engram");
|
||||
el_val_t route_similarity(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t a = query_param(path, EL_STR("a"));
|
||||
el_val_t b = query_param(path, EL_STR("b"));
|
||||
if (str_eq(a, EL_STR(""))) {
|
||||
return err_json(EL_STR("missing a"));
|
||||
}
|
||||
el_val_t snap_path = el_str_concat(dir, EL_STR("/sync-export.json"));
|
||||
engram_save(snap_path);
|
||||
el_val_t snap = fs_read(snap_path);
|
||||
if (str_eq(snap, EL_STR(""))) {
|
||||
return EL_STR("{\"nodes\":[],\"edges\":[]}");
|
||||
if (str_eq(b, EL_STR(""))) {
|
||||
return err_json(EL_STR("missing b"));
|
||||
}
|
||||
return snap;
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_save(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t p = json_get_string(body, EL_STR("path"));
|
||||
if (str_eq(p, EL_STR(""))) {
|
||||
el_val_t dir = env(EL_STR("ENGRAM_DATA_DIR"));
|
||||
if (str_eq(dir, EL_STR(""))) {
|
||||
dir = EL_STR("/tmp/engram");
|
||||
}
|
||||
p = el_str_concat(dir, EL_STR("/snapshot.json"));
|
||||
}
|
||||
engram_save(p);
|
||||
return el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"path\":\""), p), EL_STR("\"}"));
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_load(el_val_t method, el_val_t path, el_val_t body) {
|
||||
el_val_t p = json_get_string(body, EL_STR("path"));
|
||||
if (str_eq(p, EL_STR(""))) {
|
||||
el_val_t dir = env(EL_STR("ENGRAM_DATA_DIR"));
|
||||
if (str_eq(dir, EL_STR(""))) {
|
||||
dir = EL_STR("/tmp/engram");
|
||||
}
|
||||
p = el_str_concat(dir, EL_STR("/snapshot.json"));
|
||||
}
|
||||
engram_load(p);
|
||||
return ok_json();
|
||||
return 0;
|
||||
}
|
||||
|
||||
el_val_t route_health(el_val_t method, el_val_t path, el_val_t body) {
|
||||
return EL_STR("{\"status\":\"ok\",\"engine\":\"engram-runtime-native\"}");
|
||||
el_val_t sim = engram_cosine_sim(a, b);
|
||||
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"a\":\""), a), EL_STR("\",\"b\":\"")), b), EL_STR("\",\"cosine\":")), float_to_str(sim)), EL_STR("}"));
|
||||
return 0;
|
||||
}
|
||||
|
||||
@@ -329,15 +452,24 @@ el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body) {
|
||||
return route_health(method, path, body);
|
||||
}
|
||||
}
|
||||
if (str_eq(method, EL_STR("POST")) && str_starts_with(clean, EL_STR("/api/neuron/state-events"))) {
|
||||
return route_create_ise(method, path, body);
|
||||
if (str_eq(method, EL_STR("POST")) && str_eq(clean, EL_STR("/api/neuron/state-events"))) {
|
||||
return route_emit_ise(method, path, body);
|
||||
}
|
||||
if (!check_auth_ok(method, body)) {
|
||||
return err_json(EL_STR("unauthorized"));
|
||||
}
|
||||
if (str_eq(method, EL_STR("POST")) && str_eq(clean, EL_STR("/api/neuron/knowledge/capture"))) {
|
||||
return route_capture_knowledge(method, path, body);
|
||||
}
|
||||
if (str_eq(method, EL_STR("GET")) && (str_eq(clean, EL_STR("/api/stats")) || str_eq(clean, EL_STR("/stats")))) {
|
||||
return route_stats(method, path, body);
|
||||
}
|
||||
if (str_eq(method, EL_STR("GET")) && (str_eq(clean, EL_STR("/api/act-stats")) || str_eq(clean, EL_STR("/act-stats")))) {
|
||||
return route_act_stats(method, path, body);
|
||||
}
|
||||
if (str_eq(method, EL_STR("GET")) && (str_eq(clean, EL_STR("/api/text-health")) || str_eq(clean, EL_STR("/text-health")))) {
|
||||
return route_text_health(method, path, body);
|
||||
}
|
||||
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/nodes")) || str_eq(clean, EL_STR("/nodes")))) {
|
||||
return route_create_node(method, path, body);
|
||||
}
|
||||
@@ -356,6 +488,9 @@ el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body) {
|
||||
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/edges")) || str_eq(clean, EL_STR("/edges")))) {
|
||||
return route_create_edge(method, path, body);
|
||||
}
|
||||
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/edges/batch")) || str_eq(clean, EL_STR("/edges/batch")))) {
|
||||
return route_create_edges_batch(method, path, body);
|
||||
}
|
||||
if (str_eq(method, EL_STR("GET")) && str_starts_with(clean, EL_STR("/api/neighbors/"))) {
|
||||
return route_neighbors(method, path, body);
|
||||
}
|
||||
@@ -374,32 +509,46 @@ el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body) {
|
||||
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/strengthen")) || str_eq(clean, EL_STR("/strengthen")))) {
|
||||
return route_strengthen(method, path, body);
|
||||
}
|
||||
if (str_eq(method, EL_STR("GET")) && (str_eq(clean, EL_STR("/api/sync")) || str_eq(clean, EL_STR("/sync")))) {
|
||||
return route_sync(method, path, body);
|
||||
}
|
||||
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/save")) || str_eq(clean, EL_STR("/save")))) {
|
||||
return route_save(method, path, body);
|
||||
}
|
||||
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/load")) || str_eq(clean, EL_STR("/load")))) {
|
||||
return route_load(method, path, body);
|
||||
}
|
||||
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/load-merge")) || str_eq(clean, EL_STR("/load-merge")))) {
|
||||
return route_load_merge(method, path, body);
|
||||
}
|
||||
if (str_eq(method, EL_STR("GET")) && str_eq(clean, EL_STR("/api/sync"))) {
|
||||
return route_sync(method, path, body);
|
||||
}
|
||||
if (str_eq(clean, EL_STR("/api/embed-backfill"))) {
|
||||
return route_embed_backfill(method, path, body);
|
||||
}
|
||||
if (str_eq(method, EL_STR("GET")) && str_starts_with(clean, EL_STR("/api/similarity"))) {
|
||||
return route_similarity(method, path, body);
|
||||
}
|
||||
return el_str_concat(el_str_concat(EL_STR("{\"error\":\"not found\",\"path\":\""), clean), EL_STR("\"}"));
|
||||
return 0;
|
||||
}
|
||||
|
||||
int main(int _argc, char** _argv) {
|
||||
el_runtime_init_args(_argc, _argv);
|
||||
bind_str = env(EL_STR("ENGRAM_BIND"));
|
||||
if (str_eq(bind_str, EL_STR(""))) {
|
||||
bind_str = EL_STR(":8742");
|
||||
}
|
||||
bind_raw = env(EL_STR("ENGRAM_BIND"));
|
||||
bind_str = ({ el_val_t _if_result_34 = 0; if (str_eq(bind_raw, EL_STR(""))) { _if_result_34 = (EL_STR(":8742")); } else { _if_result_34 = (bind_raw); } _if_result_34; });
|
||||
port = parse_port(bind_str);
|
||||
data_dir = env(EL_STR("ENGRAM_DATA_DIR"));
|
||||
if (str_eq(data_dir, EL_STR(""))) {
|
||||
data_dir = EL_STR("/tmp/engram");
|
||||
}
|
||||
data_dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
|
||||
data_dir = ({ el_val_t _if_result_35 = 0; if (str_eq(data_dir_raw, EL_STR(""))) { _if_result_35 = (EL_STR("/tmp/engram")); } else { _if_result_35 = (data_dir_raw); } _if_result_35; });
|
||||
snapshot_path = el_str_concat(data_dir, EL_STR("/snapshot.json"));
|
||||
engram_load(snapshot_path);
|
||||
boot_snap = fs_read(snapshot_path);
|
||||
if (!str_eq(boot_snap, EL_STR(""))) {
|
||||
if (engram_node_count() == 0) {
|
||||
println(EL_STR("[engram] WARNING: snapshot.json is non-empty but load produced 0 nodes \xe2\x80\x94 preserving copy at snapshot.failed-load.json"));
|
||||
fs_write(el_str_concat(data_dir, EL_STR("/snapshot.failed-load.json")), boot_snap);
|
||||
} else {
|
||||
fs_write(el_str_concat(data_dir, EL_STR("/snapshot.boot-backup.json")), boot_snap);
|
||||
}
|
||||
}
|
||||
println(EL_STR("[engram] runtime-native graph engine"));
|
||||
println(el_str_concat(EL_STR("[engram] data_dir="), data_dir));
|
||||
println(el_str_concat(EL_STR("[engram] node_count="), int_to_str(engram_node_count())));
|
||||
|
||||
+424
-67
@@ -50,12 +50,8 @@ fn query_param(path: String, key: String) -> String {
|
||||
if pos < 0 { return "" }
|
||||
let after: String = str_slice(qs, pos + str_len(needle), str_len(qs))
|
||||
let amp: Int = str_index_of(after, "&")
|
||||
// SPEC-SEARCH-UPGRADE 2026-07-14: URL-decode the extracted value (%XX and
|
||||
// '+' were previously passed through literally, so an encoded multi-word
|
||||
// query arrived as junk tokens — pre-existing GET-path defect, masked
|
||||
// until search could actually rank multi-word queries).
|
||||
if amp < 0 { return url_decode(after) }
|
||||
url_decode(str_slice(after, 0, amp))
|
||||
if amp < 0 { return after }
|
||||
str_slice(after, 0, amp)
|
||||
}
|
||||
|
||||
fn query_int(path: String, key: String, default_val: Int) -> Int {
|
||||
@@ -80,13 +76,112 @@ fn route_stats(method: String, path: String, body: String) -> String {
|
||||
engram_stats_json()
|
||||
}
|
||||
|
||||
// route_act_stats — GET /api/act-stats
|
||||
// (2026-08-04 self-review) engram_act_stats_json() has existed since the
|
||||
// 2026-07-27 review but was reachable ONLY through the soul daemon's heartbeat
|
||||
// binding. Every activation-layer gauge — WM evictions, breakthroughs, embedder
|
||||
// breaker state, context drift, and now the Hebbian counters — was therefore
|
||||
// invisible unless the soul happened to be running and its ISEs were read back
|
||||
// out of the store. Diagnosing the activation layer required a working soul,
|
||||
// which is exactly backwards: the lower layer should be observable on its own.
|
||||
// This review needed it to verify link formation and could not get at it. One
|
||||
// line of plumbing, and the whole activation layer becomes directly diagnosable.
|
||||
fn route_act_stats(method: String, path: String, body: String) -> String {
|
||||
engram_act_stats_json()
|
||||
}
|
||||
|
||||
// route_text_health — GET /api/text-health
|
||||
// (2026-08-08 self-review) The daily census half of the text-integrity gauge.
|
||||
// Today's review found that the JSON parser had been replacing every \uXXXX
|
||||
// escape with a literal '?' for at least two months: 3,119 of 4,081
|
||||
// non-telemetry nodes (76%) were damaged, including the self traversal root
|
||||
// and every values node, and NOTHING detected it — because every gauge in the
|
||||
// system measured whether the machinery was running, and none measured whether
|
||||
// the text it carried was intact. No snapshot on disk predates the damage, so
|
||||
// it cannot be undone; it can only be made impossible to repeat quietly.
|
||||
//
|
||||
// The parser is fixed. This route is the standing check: `damaged` should now
|
||||
// hold flat at its historical floor and never climb. `write_damaged` (also on
|
||||
// the heartbeat as txt_damaged) is the live regression signal — non-zero means
|
||||
// a write path is mangling text right now.
|
||||
fn route_text_health(method: String, path: String, body: String) -> String {
|
||||
engram_text_health_json()
|
||||
}
|
||||
|
||||
// (2026-07-18 self-review) Scoping sweep: `let` inside an if-block creates an
|
||||
// inner scope only — it does NOT mutate the outer binding (documented with
|
||||
// evidence in awareness.el, 2026-05-25). Every default/reassignment below used
|
||||
// that broken pattern, so defaults never applied: nodes were created with
|
||||
// node_type="" and salience=0.0, /api/search and /api/activate ALWAYS ran with
|
||||
// q="" regardless of input, edges defaulted to relation=""/weight=0.0, and
|
||||
// save/load with no "path" hit engram_save(""). Rewritten to the
|
||||
// `let x = if cond { a } else { b }` expression form (the pattern the newer
|
||||
// routes route_emit_ise/route_capture_knowledge already use correctly).
|
||||
// persist_canonical — save the canonical snapshot after a durable write.
|
||||
//
|
||||
// WHY (2026-07-22 self-review): the 2026-07-21 fix correctly stopped READ
|
||||
// routes from writing the canonical snapshot.json — but nothing was left
|
||||
// that saved it on WRITE. Every mutation (node create, edge create,
|
||||
// knowledge capture, forget, merge) lived only in RAM until someone POSTed
|
||||
// /api/save manually; a process restart silently discarded everything since
|
||||
// the last manual save. Observed live: two engram restarts during the
|
||||
// 2026-07-22 review reverted the store to a ~17h-old snapshot, destroying
|
||||
// same-day writes. Reads must never write the canonical; writes must always
|
||||
// persist it. ISE telemetry is deliberately excluded (48h-pruned, loss-
|
||||
// tolerant, ~2/min — snapshotting the whole store per heartbeat is waste;
|
||||
// any durable write that follows persists the pruning too).
|
||||
fn persist_canonical() -> Int {
|
||||
let dir_raw: String = env("ENGRAM_DATA_DIR")
|
||||
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
|
||||
// (2026-08-10 self-review) This returned a hardcoded 1, which made every
|
||||
// caller's `let saved: Int = persist_canonical()` a dead variable — six
|
||||
// durable write paths each believed they had confirmation of a successful
|
||||
// canonical persist and none of them had any. Propagate the real result.
|
||||
return engram_save(dir + "/snapshot.json")
|
||||
}
|
||||
|
||||
// INCOMPLETE-ROUTE FIX (2026-07-24 self-review): this route silently dropped
|
||||
// label, importance, tier, and tags — engram_node() defaults label to content
|
||||
// and importance to 0.5, so every node created over HTTP lost its metadata.
|
||||
// Observed live: the soul's boot-counter write-back landed with
|
||||
// label="soul:boot_count:99" (content), importance 0.5, no tags. Honor the
|
||||
// full field set via engram_node_full when any of them is supplied.
|
||||
// PRESENCE-AWARE DEFAULTS (2026-08-01 self-review): the old pattern
|
||||
// `if x == 0.0 { default }` made a legitimate 0.0 unrepresentable — a caller
|
||||
// setting salience/importance/weight to zero silently got 0.5. json_get_raw
|
||||
// returns "" when the key is ABSENT and the raw token when present, so
|
||||
// absence and zero are now distinguishable. Also: confidence was hardcoded
|
||||
// to 1.0 regardless of input — every HTTP-created node claimed full
|
||||
// epistemic confidence. Now honored from the payload (default 1.0).
|
||||
fn route_create_node(method: String, path: String, body: String) -> String {
|
||||
let content: String = json_get_string(body, "content")
|
||||
let node_type: String = json_get_string(body, "node_type")
|
||||
if str_eq(node_type, "") { let node_type = "Memory" }
|
||||
let salience: Float = json_get_float(body, "salience")
|
||||
if salience == 0.0 { let salience = 0.5 }
|
||||
let id: String = engram_node(content, node_type, salience)
|
||||
let nt_raw: String = json_get_string(body, "node_type")
|
||||
let node_type: String = if str_eq(nt_raw, "") { "Memory" } else { nt_raw }
|
||||
let sal_present: String = json_get_raw(body, "salience")
|
||||
let salience: Float = if str_eq(sal_present, "") { 0.5 } else { json_get_float(body, "salience") }
|
||||
let label_raw: String = json_get_string(body, "label")
|
||||
let label: String = if str_eq(label_raw, "") { content } else { label_raw }
|
||||
let imp_present: String = json_get_raw(body, "importance")
|
||||
let importance: Float = if str_eq(imp_present, "") { 0.5 } else { json_get_float(body, "importance") }
|
||||
let conf_present: String = json_get_raw(body, "confidence")
|
||||
let confidence: Float = if str_eq(conf_present, "") { 1.0 } else { json_get_float(body, "confidence") }
|
||||
let tier_raw: String = json_get_string(body, "tier")
|
||||
let tier: String = if str_eq(tier_raw, "") { "Working" } else { tier_raw }
|
||||
let tags: String = json_get_string(body, "tags")
|
||||
// NO el_from_float WRAPPER (2026-08-01 self-review): salience/importance/
|
||||
// confidence are already Float (el_val_t) values — json_get_float and
|
||||
// Float literals both encode. Wrapping them in el_from_float AGAIN
|
||||
// reinterpreted the boxed bits as a raw double, producing garbage that
|
||||
// failed engram_decode_score's range check and clamped every HTTP-created
|
||||
// node to defaults (salience 0.9 in → 0.5 stored; confidence 0.6 in → 1.0
|
||||
// stored — verified live). route_emit_ise always passed Floats bare and
|
||||
// its 0.3/0.3/0.8 stored correctly; this call now does the same.
|
||||
let id: String = engram_node_full(
|
||||
content, node_type, label,
|
||||
salience, importance, confidence,
|
||||
tier, tags
|
||||
)
|
||||
let saved: Int = persist_canonical()
|
||||
"{\"id\":\"" + id + "\",\"content\":\"" + content + "\",\"node_type\":\"" + node_type + "\"}"
|
||||
}
|
||||
|
||||
@@ -107,13 +202,14 @@ fn route_scan_nodes(method: String, path: String, body: String) -> String {
|
||||
}
|
||||
|
||||
// route_scan_edges — bulk export of all edges as a JSON array. Implemented
|
||||
// via engram_save → fs_read of the canonical on-disk snapshot, which the
|
||||
// runtime keeps in lockstep with the in-memory graph. Live against the
|
||||
// running graph, not a stale export.
|
||||
// via engram_save → fs_read of a SCRATCH export path. (2026-07-21 self-review:
|
||||
// previously this saved over the canonical snapshot.json on every GET — if the
|
||||
// process ever booted with a partial/empty store, the first read request
|
||||
// clobbered the good snapshot. Read routes must never write the canonical path.)
|
||||
fn route_scan_edges(method: String, path: String, body: String) -> String {
|
||||
let dir: String = env("ENGRAM_DATA_DIR")
|
||||
if str_eq(dir, "") { let dir = "/tmp/engram" }
|
||||
let snap_path: String = dir + "/snapshot.json"
|
||||
let dir_raw: String = env("ENGRAM_DATA_DIR")
|
||||
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
|
||||
let snap_path: String = dir + "/.scan-export.json"
|
||||
engram_save(snap_path)
|
||||
let snap: String = fs_read(snap_path)
|
||||
if str_eq(snap, "") { return "[]" }
|
||||
@@ -126,43 +222,88 @@ fn route_scan_edges(method: String, path: String, body: String) -> String {
|
||||
}
|
||||
|
||||
fn route_search(method: String, path: String, body: String) -> String {
|
||||
let q: String = ""
|
||||
if str_eq(method, "GET") {
|
||||
let q = query_param(path, "q")
|
||||
} else {
|
||||
let q = json_get_string(body, "query")
|
||||
}
|
||||
let limit: Int = query_int(path, "limit", 20)
|
||||
if limit == 0 { let limit = json_get_int(body, "limit") }
|
||||
if limit == 0 { let limit = 20 }
|
||||
let q: String = if str_eq(method, "GET") { query_param(path, "q") } else { json_get_string(body, "query") }
|
||||
let lim_url: Int = query_int(path, "limit", 0)
|
||||
let lim_body: Int = json_get_int(body, "limit")
|
||||
let lim_either: Int = if lim_url > 0 { lim_url } else { lim_body }
|
||||
let limit: Int = if lim_either > 0 { lim_either } else { 20 }
|
||||
return engram_search_json(q, limit)
|
||||
}
|
||||
|
||||
fn route_activate(method: String, path: String, body: String) -> String {
|
||||
let q: String = ""
|
||||
let depth: Int = 3
|
||||
if str_eq(method, "GET") {
|
||||
let q = query_param(path, "q")
|
||||
let depth = query_int(path, "depth", 3)
|
||||
} else {
|
||||
let q = json_get_string(body, "query")
|
||||
let bd: Int = json_get_int(body, "depth")
|
||||
if bd > 0 { let depth = bd }
|
||||
}
|
||||
let q: String = if str_eq(method, "GET") { query_param(path, "q") } else { json_get_string(body, "query") }
|
||||
// Guard: engram_activate with an empty query matches zero seeds, which
|
||||
// zeroes ALL carried working-memory weights (documented in awareness.el
|
||||
// perceive()). Never let an empty activation through to wipe WM.
|
||||
if str_eq(q, "") { return err_json("missing query") }
|
||||
let d_raw: Int = if str_eq(method, "GET") { query_int(path, "depth", 3) } else { json_get_int(body, "depth") }
|
||||
let depth: Int = if d_raw > 0 { d_raw } else { 3 }
|
||||
return "{\"results\":" + engram_activate_json(q, depth) + "}"
|
||||
}
|
||||
|
||||
fn route_create_edge(method: String, path: String, body: String) -> String {
|
||||
let from_id: String = json_get_string(body, "from_id")
|
||||
let to_id: String = json_get_string(body, "to_id")
|
||||
let relation: String = json_get_string(body, "relation")
|
||||
if str_eq(relation, "") { let relation = "associates" }
|
||||
let weight: Float = json_get_float(body, "weight")
|
||||
if weight == 0.0 { let weight = 0.5 }
|
||||
let rel_raw: String = json_get_string(body, "relation")
|
||||
let relation: String = if str_eq(rel_raw, "") { "associates" } else { rel_raw }
|
||||
// Presence-aware (2026-08-01): weight 0.0 is a legitimate edge weight
|
||||
// (dormant association); only default when the key is absent.
|
||||
let w_present: String = json_get_raw(body, "weight")
|
||||
let weight: Float = if str_eq(w_present, "") { 0.5 } else { json_get_float(body, "weight") }
|
||||
engram_connect(from_id, to_id, weight, relation)
|
||||
let saved: Int = persist_canonical()
|
||||
"{\"ok\":true,\"from_id\":\"" + from_id + "\",\"to_id\":\"" + to_id + "\",\"relation\":\"" + relation + "\"}"
|
||||
}
|
||||
|
||||
// route_create_edges_batch — POST /api/edges/batch {"edges":[{from_id,to_id,relation,weight}, ...]}
|
||||
//
|
||||
// WHY THIS EXISTS (2026-08-07 self-review). persist_canonical() writes the
|
||||
// FULL canonical snapshot — 60MB at current graph size — and route_create_edge
|
||||
// calls it once per edge. That is correct for the interactive one-edge case and
|
||||
// ruinous for any bulk write: the soul's Hebbian consolidation path delivers
|
||||
// ~14 associations per 8-minute heartbeat, which through the single-edge route
|
||||
// would be ~840MB of disk writes per beat, ~150GB/day, to persist 14 edges.
|
||||
//
|
||||
// The fix is not to weaken durability — it is to make the unit of durability
|
||||
// the BATCH. Connect every edge, then snapshot exactly once. Same guarantee
|
||||
// (nothing acknowledged is lost to a restart), 1/N the writes. Empty or
|
||||
// malformed entries are skipped rather than aborting the batch: a consolidation
|
||||
// payload is best-effort by design, and one bad id should not cost the other 13.
|
||||
//
|
||||
// Returns the accepted count so the caller can tell delivery from silence.
|
||||
fn route_create_edges_batch(method: String, path: String, body: String) -> String {
|
||||
let arr: String = json_get_raw(body, "edges")
|
||||
if str_eq(arr, "") { return err_json("missing edges array") }
|
||||
let n: Int = json_array_len(arr)
|
||||
if n == 0 { return "{\"ok\":true,\"accepted\":0,\"skipped\":0}" }
|
||||
let i: Int = 0
|
||||
let accepted: Int = 0
|
||||
let skipped: Int = 0
|
||||
while i < n {
|
||||
let item: String = json_array_get(arr, i)
|
||||
let from_id: String = json_get_string(item, "from_id")
|
||||
let to_id: String = json_get_string(item, "to_id")
|
||||
if str_eq(from_id, "") || str_eq(to_id, "") {
|
||||
let skipped = skipped + 1
|
||||
} else {
|
||||
let rel_raw: String = json_get_string(item, "relation")
|
||||
let relation: String = if str_eq(rel_raw, "") { "associates" } else { rel_raw }
|
||||
let w_present: String = json_get_raw(item, "weight")
|
||||
let weight: Float = if str_eq(w_present, "") { 0.5 } else { json_get_float(item, "weight") }
|
||||
engram_connect(from_id, to_id, weight, relation)
|
||||
let accepted = accepted + 1
|
||||
}
|
||||
let i = i + 1
|
||||
}
|
||||
// ONE snapshot for the whole batch — the entire point of this route.
|
||||
// Skip it when nothing was accepted: an all-malformed payload must not
|
||||
// trigger a 60MB write.
|
||||
if accepted > 0 {
|
||||
let saved: Int = persist_canonical()
|
||||
}
|
||||
return "{\"ok\":true,\"accepted\":" + int_to_str(accepted) + ",\"skipped\":" + int_to_str(skipped) + "}"
|
||||
}
|
||||
|
||||
fn route_neighbors(method: String, path: String, body: String) -> String {
|
||||
let id: String = extract_id(path, "/api/neighbors/")
|
||||
if str_eq(id, "") { return err_json("missing id") }
|
||||
@@ -174,6 +315,7 @@ fn route_strengthen(method: String, path: String, body: String) -> String {
|
||||
let id: String = json_get_string(body, "node_id")
|
||||
if str_eq(id, "") { return err_json("missing node_id") }
|
||||
engram_strengthen(id)
|
||||
let saved: Int = persist_canonical()
|
||||
ok_json()
|
||||
}
|
||||
|
||||
@@ -181,33 +323,84 @@ fn route_forget(method: String, path: String, body: String) -> String {
|
||||
let id: String = extract_id(path, "/api/nodes/")
|
||||
if str_eq(id, "") { return err_json("missing id") }
|
||||
engram_forget(id)
|
||||
let saved: Int = persist_canonical()
|
||||
ok_json()
|
||||
}
|
||||
|
||||
fn route_save(method: String, path: String, body: String) -> String {
|
||||
let p: String = json_get_string(body, "path")
|
||||
if str_eq(p, "") {
|
||||
let dir: String = env("ENGRAM_DATA_DIR")
|
||||
if str_eq(dir, "") { let dir = "/tmp/engram" }
|
||||
let p = dir + "/snapshot.json"
|
||||
}
|
||||
engram_save(p)
|
||||
"{\"ok\":true,\"path\":\"" + p + "\"}"
|
||||
let p_raw: String = json_get_string(body, "path")
|
||||
let dir_raw: String = env("ENGRAM_DATA_DIR")
|
||||
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
|
||||
let p: String = if str_eq(p_raw, "") { dir + "/snapshot.json" } else { p_raw }
|
||||
// (2026-08-10 self-review) engram_save returns 0 on an empty path and the
|
||||
// route discarded it, so the response was a literal "ok":true regardless
|
||||
// of whether anything was written. Report the actual result AND the counts
|
||||
// that were supposed to have been written — the same move that made
|
||||
// route_health honest on 2026-08-01. A caller can now tell "saved 13k
|
||||
// nodes" from "saved nothing and said ok".
|
||||
let sv: Int = engram_save(p)
|
||||
let sv_ok: String = if sv == 0 { "false" } else { "true" }
|
||||
"{\"ok\":" + sv_ok + ",\"path\":\"" + p + "\",\"node_count\":" + int_to_str(engram_node_count()) + ",\"edge_count\":" + int_to_str(engram_edge_count()) + "}"
|
||||
}
|
||||
|
||||
fn route_load(method: String, path: String, body: String) -> String {
|
||||
let p: String = json_get_string(body, "path")
|
||||
if str_eq(p, "") {
|
||||
let dir: String = env("ENGRAM_DATA_DIR")
|
||||
if str_eq(dir, "") { let dir = "/tmp/engram" }
|
||||
let p = dir + "/snapshot.json"
|
||||
}
|
||||
engram_load(p)
|
||||
ok_json()
|
||||
let p_raw: String = json_get_string(body, "path")
|
||||
let dir_raw: String = env("ENGRAM_DATA_DIR")
|
||||
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
|
||||
let p: String = if str_eq(p_raw, "") { dir + "/snapshot.json" } else { p_raw }
|
||||
// (2026-08-10 self-review) This was a stub response over the single most
|
||||
// destructive operation in the server. engram_load returns 0 on an empty
|
||||
// path, an unopenable file, a zero-length file, or malloc failure — and
|
||||
// this route answered ok_json() in every one of those cases.
|
||||
//
|
||||
// Precise failure shape (el_runtime.c:9890): the fopen guard runs BEFORE
|
||||
// the store reset, so a MISSING path is genuinely safe — it returns 0 with
|
||||
// the graph intact. The dangerous case is a readable-but-malformed file:
|
||||
// the reset loop frees every node and edge FIRST, then parses, so a
|
||||
// truncated or non-snapshot JSON leaves a hollow store — and the caller
|
||||
// was told "ok":true. With 37 GB of stale dated snapshots sitting in the
|
||||
// data dir as tempting restore targets, "restore reported success and
|
||||
// silently emptied the graph" is a live risk, not a hypothetical one.
|
||||
//
|
||||
// Fix: surface the return value AND the resulting counts. node_count=0
|
||||
// after a load is the unambiguous hollow-store signal (same convention
|
||||
// route_health adopted 2026-08-01). Callers can now verify a restore
|
||||
// instead of trusting it.
|
||||
let ld: Int = engram_load(p)
|
||||
let ld_ok: String = if ld == 0 { "false" } else { "true" }
|
||||
let nc_after: Int = engram_node_count()
|
||||
let hollow: String = if nc_after == 0 { "true" } else { "false" }
|
||||
"{\"ok\":" + ld_ok + ",\"path\":\"" + p + "\",\"node_count\":" + int_to_str(nc_after) + ",\"edge_count\":" + int_to_str(engram_edge_count()) + ",\"hollow\":" + hollow + "}"
|
||||
}
|
||||
|
||||
// (2026-08-01 self-review) Health previously returned a hardcoded literal —
|
||||
// it reported "ok" even when the snapshot failed to load and the store was
|
||||
// empty. Now reports live counts so a monitor can distinguish "up and
|
||||
// loaded" from "up and hollow" (node_count=0 after boot = failed load).
|
||||
fn route_health(method: String, path: String, body: String) -> String {
|
||||
"{\"status\":\"ok\",\"engine\":\"engram-runtime-native\"}"
|
||||
"{\"status\":\"ok\",\"engine\":\"engram-runtime-native\",\"node_count\":" + int_to_str(engram_node_count()) + ",\"edge_count\":" + int_to_str(engram_edge_count()) + "}"
|
||||
}
|
||||
|
||||
// route_embed_backfill — GET/POST /api/embed-backfill?n=48
|
||||
//
|
||||
// (2026-07-25 self-review) The lazy embedding backfill runs only inside
|
||||
// engram_activate, and nothing in production calls /api/activate on this
|
||||
// store — the soul's curiosity loop activates its own in-process graph.
|
||||
// After a restart from a snapshot without vectors, embedded_count stalled
|
||||
// at 93/12175 and would never recover. This route lets the soul's
|
||||
// heartbeat pump the backfill explicitly (48/min clears a 12k backlog in
|
||||
// ~4h). Persists the canonical snapshot whenever new vectors were
|
||||
// generated — the 2026-07-25 regression happened precisely because 3747
|
||||
// in-RAM embeddings were never snapshotted before a restart. Self-
|
||||
// limiting: once coverage is full, embedded=0 and no save occurs.
|
||||
fn route_embed_backfill(method: String, path: String, body: String) -> String {
|
||||
let n: Int = query_int(path, "n", 32)
|
||||
let result: String = engram_embed_backfill(n)
|
||||
let done: Float = json_get_float(result, "embedded")
|
||||
if done > 0.0 {
|
||||
let saved: Int = persist_canonical()
|
||||
}
|
||||
return result
|
||||
}
|
||||
|
||||
// route_sync — return a snapshot of non-ISE/non-Working nodes for the soul daemon
|
||||
@@ -223,15 +416,45 @@ fn route_health(method: String, path: String, body: String) -> String {
|
||||
// (it skips nodes already present by ID). Auth-exempt: same-host internal call.
|
||||
// (2026-06-27 self-review: added this route to fix silent 10-min sync failures)
|
||||
fn route_sync(method: String, path: String, body: String) -> String {
|
||||
let dir: String = env("ENGRAM_DATA_DIR")
|
||||
if str_eq(dir, "") { let dir = "/tmp/engram" }
|
||||
let snap_path: String = dir + "/snapshot.json"
|
||||
let dir_raw: String = env("ENGRAM_DATA_DIR")
|
||||
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
|
||||
// 2026-07-21 self-review: export to a scratch path, never the canonical
|
||||
// snapshot.json — read routes must not be able to clobber the good snapshot.
|
||||
let snap_path: String = dir + "/.sync-export.json"
|
||||
engram_save(snap_path)
|
||||
let snap: String = fs_read(snap_path)
|
||||
if str_eq(snap, "") { return "{\"nodes\":[],\"edges\":[]}" }
|
||||
// 2026-08-02 self-review: this used to return {"nodes":[],"edges":[]} when
|
||||
// the export/read failed. The soul's sync_ok test (awareness.el) only
|
||||
// checks for "" and "{}", so that placeholder PASSED as a healthy sync:
|
||||
// soul.last_sync_ok_ts got stamped, sync_age_ms stayed green, the
|
||||
// sync_empty warn ISE never fired, and engram_sync reported added:0
|
||||
// forever. A totally broken sync was indistinguishable from a quiet
|
||||
// healthy one — the exact failure class this route was added to fix in
|
||||
// the first place (see 2026-06-27 note above). Return a real error so the
|
||||
// failure is loud on both sides.
|
||||
if str_eq(snap, "") { return err_json("sync export failed: snapshot unreadable") }
|
||||
return snap
|
||||
}
|
||||
|
||||
// route_load_merge — POST /api/load-merge {"path": "..."} — merge a snapshot
|
||||
// file into the live store WITHOUT resetting it (engram_load_merge skips nodes
|
||||
// already present by id). Added 2026-07-21 self-review to restore the 244 kn-
|
||||
// identity Knowledge nodes lost from the snapshot lineage between 05-13 and
|
||||
// 07-13. Requires an explicit path: refuses to run without one so it can never
|
||||
// be triggered accidentally against a default.
|
||||
fn route_load_merge(method: String, path: String, body: String) -> String {
|
||||
let p: String = json_get_string(body, "path")
|
||||
if str_eq(p, "") { return err_json("path is required") }
|
||||
if str_eq(fs_read(p), "") { return err_json("file missing or empty") }
|
||||
let before_n: Int = engram_node_count()
|
||||
let before_e: Int = engram_edge_count()
|
||||
engram_load_merge(p)
|
||||
let added_n: Int = engram_node_count() - before_n
|
||||
let added_e: Int = engram_edge_count() - before_e
|
||||
let saved: Int = persist_canonical()
|
||||
"{\"ok\":true,\"nodes_added\":" + int_to_str(added_n) + ",\"edges_added\":" + int_to_str(added_e) + ",\"node_count\":" + int_to_str(engram_node_count()) + "}"
|
||||
}
|
||||
|
||||
// route_emit_ise — write an InternalStateEvent node from the soul daemon.
|
||||
//
|
||||
// Endpoint: POST /api/neuron/state-events
|
||||
@@ -245,10 +468,20 @@ fn route_sync(method: String, path: String, body: String) -> String {
|
||||
//
|
||||
// Salience/importance set to match engram_node_full ISE defaults used by the
|
||||
// in-process fallback path in awareness.el (salience=0.3, importance=0.3,
|
||||
// confidence=0.8, tier=Episodic). High temporal_decay_rate (1.617) — ISEs
|
||||
// are inherently transient; they should decay faster than structural knowledge.
|
||||
// confidence=0.8, tier=Episodic).
|
||||
// (2026-06-26 self-review: added this route after discovering ise_post was
|
||||
// silently failing — the soul posts here but the endpoint didn't exist.)
|
||||
//
|
||||
// Retention (2026-07-16 self-review): an earlier comment here claimed ISEs
|
||||
// got temporal_decay_rate=1.617 — that was never implemented (engram_node_full
|
||||
// hardcodes 0.0), and per-node decay only dampens activation anyway; it never
|
||||
// removes nodes. By 2026-07-16 ISEs were 75% of the store (10,175 of 13,522
|
||||
// nodes, ~4,300/day, unbounded). ISEs are already WM-excluded in
|
||||
// engram_activate, so the fix is retention, not decay: every insert calls
|
||||
// engram_prune_telemetry(), a single O(nodes+edges) compaction pass that
|
||||
// removes ISEs older than ENGRAM_ISE_RETENTION_MS (default 48h), protecting
|
||||
// "session-start" labels and self_review events as durable history. At
|
||||
// ~3 ISEs/min this bounds telemetry at ~8.6k nodes instead of growing forever.
|
||||
fn route_emit_ise(method: String, path: String, body: String) -> String {
|
||||
let content: String = json_get_string(body, "content")
|
||||
if str_eq(content, "") { return err_json("missing content") }
|
||||
@@ -260,9 +493,86 @@ fn route_emit_ise(method: String, path: String, body: String) -> String {
|
||||
sal, imp, conf,
|
||||
"Episodic", "[\"internal-state\",\"InternalStateEvent\"]"
|
||||
)
|
||||
let ret_raw: String = env("ENGRAM_ISE_RETENTION_MS")
|
||||
let ret_ms: Int = if str_eq(ret_raw, "") { 172800000 } else { str_to_int(ret_raw) }
|
||||
let pruned: Int = engram_prune_telemetry(ret_ms)
|
||||
"{\"ok\":true,\"id\":\"" + id + "\",\"pruned\":" + int_to_str(pruned) + "}"
|
||||
}
|
||||
|
||||
// ── Knowledge capture ─────────────────────────────────────────────────────────
|
||||
//
|
||||
// route_capture_knowledge — direct Knowledge-node capture over HTTP.
|
||||
//
|
||||
// Endpoint: POST /api/neuron/knowledge/capture (auth required: "_auth" in body)
|
||||
// Body: {"content": "...", "title": "...", "category": "...",
|
||||
// "tier": "note|lesson|canonical", "tags": [...], "project": "...",
|
||||
// "_auth": "<key>"}
|
||||
//
|
||||
// WHY (2026-07-15 self-review): the world-ingestor integrator was designed
|
||||
// against this endpoint (its MCP-unavailable fallback), but the route never
|
||||
// existed — every direct push 404'd, and because the auth gate ran before
|
||||
// routing, the failure surfaced as {"error":"unauthorized"} and was
|
||||
// misdiagnosed for two weeks while world knowledge silently dropped.
|
||||
// POST /api/nodes was no substitute: it discards label/tags/tier, which
|
||||
// makes captured knowledge invisible to tag-scoped search and curiosity.
|
||||
//
|
||||
// The incoming knowledge tier (note/lesson/canonical) is preserved as a
|
||||
// "tier:<x>" tag rather than mapped onto Engram's cognitive tiers — Knowledge
|
||||
// nodes land in Semantic (stable reference), and the epistemic tier stays
|
||||
// queryable without inventing a lossy mapping.
|
||||
fn route_capture_knowledge(method: String, path: String, body: String) -> String {
|
||||
let content: String = json_get_string(body, "content")
|
||||
if str_eq(content, "") { return err_json("missing content") }
|
||||
let title: String = json_get_string(body, "title")
|
||||
let label: String = if str_eq(title, "") { str_slice(content, 0, 60) } else { title }
|
||||
let category_raw: String = json_get_string(body, "category")
|
||||
let category: String = if str_eq(category_raw, "") { "other" } else { category_raw }
|
||||
let ktier_raw: String = json_get_string(body, "tier")
|
||||
let ktier: String = if str_eq(ktier_raw, "") { "note" } else { ktier_raw }
|
||||
let project: String = json_get_string(body, "project")
|
||||
let tags_raw: String = json_get_raw(body, "tags")
|
||||
let tags_base: String = if str_eq(tags_raw, "") { "[]" } else { tags_raw }
|
||||
// Merge category/tier/project markers into the tag array. Search matches
|
||||
// against the tags string, so these make captures findable by facet.
|
||||
let base_len: Int = str_len(tags_base)
|
||||
let head: String = str_slice(tags_base, 0, base_len - 1)
|
||||
let sep: String = if str_eq(head, "[") { "" } else { "," }
|
||||
let safe_cat: String = str_replace(category, "\"", "'")
|
||||
let safe_tier: String = str_replace(ktier, "\"", "'")
|
||||
let safe_proj: String = str_replace(project, "\"", "'")
|
||||
let proj_tag: String = if str_eq(safe_proj, "") { "" } else { ",\"project:" + safe_proj + "\"" }
|
||||
let tags: String = head + sep + "\"category:" + safe_cat + "\",\"tier:" + safe_tier + "\"" + proj_tag + "]"
|
||||
let sal: Float = 0.5
|
||||
let imp: Float = 0.5
|
||||
let conf: Float = 0.9
|
||||
let id: String = engram_node_full(
|
||||
content, "Knowledge", label,
|
||||
sal, imp, conf,
|
||||
"Semantic", tags
|
||||
)
|
||||
let saved: Int = persist_canonical()
|
||||
"{\"ok\":true,\"id\":\"" + id + "\"}"
|
||||
}
|
||||
|
||||
// route_similarity — GET /api/similarity?a=<id>&b=<id>
|
||||
//
|
||||
// (2026-08-01 self-review) engram_cosine_sim was added 2026-07-24
|
||||
// (bl-b2d1c944) with the stated purpose of exposing semantic distance to
|
||||
// "EL code and the introspection API" — but it had ZERO callers anywhere:
|
||||
// no route, no soul-daemon use. The activation path uses embeddings
|
||||
// internally (semantic seeding, Pass-2 additive term), but there was no way
|
||||
// to probe pairwise node similarity from outside. This closes that: cosine
|
||||
// in [-1,1], or -2 when either node is missing or not yet embedded (so
|
||||
// "not comparable" is distinguishable from "genuinely orthogonal" 0.0).
|
||||
fn route_similarity(method: String, path: String, body: String) -> String {
|
||||
let a: String = query_param(path, "a")
|
||||
let b: String = query_param(path, "b")
|
||||
if str_eq(a, "") { return err_json("missing a") }
|
||||
if str_eq(b, "") { return err_json("missing b") }
|
||||
let sim: Float = engram_cosine_sim(a, b)
|
||||
"{\"a\":\"" + a + "\",\"b\":\"" + b + "\",\"cosine\":" + float_to_str(sim) + "}"
|
||||
}
|
||||
|
||||
// ── Auth ──────────────────────────────────────────────────────────────────────
|
||||
|
||||
fn check_auth_ok(method: String, body: String) -> Bool {
|
||||
@@ -299,10 +609,22 @@ fn handle_request(method: String, path: String, body: String) -> String {
|
||||
return err_json("unauthorized")
|
||||
}
|
||||
|
||||
// Knowledge capture (auth enforced above; the world-ingestor integrator
|
||||
// and any headless session without MCP push knowledge through this)
|
||||
if str_eq(method, "POST") && str_eq(clean, "/api/neuron/knowledge/capture") {
|
||||
return route_capture_knowledge(method, path, body)
|
||||
}
|
||||
|
||||
// Stats
|
||||
if str_eq(method, "GET") && (str_eq(clean, "/api/stats") || str_eq(clean, "/stats")) {
|
||||
return route_stats(method, path, body)
|
||||
}
|
||||
if str_eq(method, "GET") && (str_eq(clean, "/api/act-stats") || str_eq(clean, "/act-stats")) {
|
||||
return route_act_stats(method, path, body)
|
||||
}
|
||||
if str_eq(method, "GET") && (str_eq(clean, "/api/text-health") || str_eq(clean, "/text-health")) {
|
||||
return route_text_health(method, path, body)
|
||||
}
|
||||
|
||||
// Nodes
|
||||
if str_eq(method, "POST") && (str_eq(clean, "/api/nodes") || str_eq(clean, "/nodes")) {
|
||||
@@ -325,6 +647,13 @@ fn handle_request(method: String, path: String, body: String) -> String {
|
||||
if str_eq(method, "POST") && (str_eq(clean, "/api/edges") || str_eq(clean, "/edges")) {
|
||||
return route_create_edge(method, path, body)
|
||||
}
|
||||
// Batch edge write — one snapshot for the whole payload. Must be tested
|
||||
// BEFORE nothing else claims it; the exact-match on "/api/edges" above
|
||||
// does not catch "/api/edges/batch", so order is not load-bearing here,
|
||||
// but keeping the two adjacent keeps them from drifting apart.
|
||||
if str_eq(method, "POST") && (str_eq(clean, "/api/edges/batch") || str_eq(clean, "/edges/batch")) {
|
||||
return route_create_edges_batch(method, path, body)
|
||||
}
|
||||
if str_eq(method, "GET") && str_starts_with(clean, "/api/neighbors/") {
|
||||
return route_neighbors(method, path, body)
|
||||
}
|
||||
@@ -355,27 +684,55 @@ fn handle_request(method: String, path: String, body: String) -> String {
|
||||
if str_eq(method, "POST") && (str_eq(clean, "/api/load") || str_eq(clean, "/load")) {
|
||||
return route_load(method, path, body)
|
||||
}
|
||||
if str_eq(method, "POST") && (str_eq(clean, "/api/load-merge") || str_eq(clean, "/load-merge")) {
|
||||
return route_load_merge(method, path, body)
|
||||
}
|
||||
|
||||
// Sync — soul daemon periodic pull of non-ISE knowledge into in-process graph
|
||||
if str_eq(method, "GET") && str_eq(clean, "/api/sync") {
|
||||
return route_sync(method, path, body)
|
||||
}
|
||||
|
||||
// Embedding backfill — pumped by the soul heartbeat (2026-07-25)
|
||||
if str_eq(clean, "/api/embed-backfill") {
|
||||
return route_embed_backfill(method, path, body)
|
||||
}
|
||||
|
||||
// Semantic similarity probe (2026-08-01)
|
||||
if str_eq(method, "GET") && str_starts_with(clean, "/api/similarity") {
|
||||
return route_similarity(method, path, body)
|
||||
}
|
||||
|
||||
"{\"error\":\"not found\",\"path\":\"" + clean + "\"}"
|
||||
}
|
||||
|
||||
// ── Entry ─────────────────────────────────────────────────────────────────────
|
||||
|
||||
let bind_str: String = env("ENGRAM_BIND")
|
||||
if str_eq(bind_str, "") { let bind_str = ":8742" }
|
||||
let bind_raw: String = env("ENGRAM_BIND")
|
||||
let bind_str: String = if str_eq(bind_raw, "") { ":8742" } else { bind_raw }
|
||||
let port: Int = parse_port(bind_str)
|
||||
|
||||
// On startup, try to load any existing snapshot (best effort).
|
||||
let data_dir: String = env("ENGRAM_DATA_DIR")
|
||||
if str_eq(data_dir, "") { let data_dir = "/tmp/engram" }
|
||||
let data_dir_raw: String = env("ENGRAM_DATA_DIR")
|
||||
let data_dir: String = if str_eq(data_dir_raw, "") { "/tmp/engram" } else { data_dir_raw }
|
||||
let snapshot_path: String = data_dir + "/snapshot.json"
|
||||
engram_load(snapshot_path)
|
||||
|
||||
// 2026-07-21 self-review boot guard: if the snapshot file has content but the
|
||||
// load produced 0 nodes, something is wrong (corrupt file / parse failure).
|
||||
// Preserve the evidence and warn loudly — and since read routes no longer write
|
||||
// the canonical path, a bad boot can no longer clobber the good snapshot.
|
||||
let boot_snap: String = fs_read(snapshot_path)
|
||||
if !str_eq(boot_snap, "") {
|
||||
if engram_node_count() == 0 {
|
||||
println("[engram] WARNING: snapshot.json is non-empty but load produced 0 nodes — preserving copy at snapshot.failed-load.json")
|
||||
fs_write(data_dir + "/snapshot.failed-load.json", boot_snap)
|
||||
} else {
|
||||
// Good load: keep a boot-time backup of the snapshot as loaded.
|
||||
fs_write(data_dir + "/snapshot.boot-backup.json", boot_snap)
|
||||
}
|
||||
}
|
||||
|
||||
println("[engram] runtime-native graph engine")
|
||||
println("[engram] data_dir=" + data_dir)
|
||||
println("[engram] node_count=" + int_to_str(engram_node_count()))
|
||||
|
||||
@@ -1545,6 +1545,17 @@ typedef struct {
|
||||
#endif
|
||||
} HttpWorkerArg;
|
||||
|
||||
/* Forward declarations for the loopback/API-key hardening helpers defined
|
||||
* further down. Without these, http_worker's calls below were implicit
|
||||
* declarations and the later `static` definitions conflicted with them — this
|
||||
* file did not compile at all. (2026-08-08 self-review: the hardening work
|
||||
* they belong to had been sitting uncommitted in the working tree since
|
||||
* 2026-07-15 in exactly this non-building state, which is presumably why it
|
||||
* was never committed. Adding the two prototypes is the whole fix.) */
|
||||
static int el_http_request_authorized(const char* method, const char* path,
|
||||
const char* hdr_block);
|
||||
static void el_http_send_401(int fd);
|
||||
|
||||
static void* http_worker(void* arg) {
|
||||
HttpWorkerArg* a = (HttpWorkerArg*)arg;
|
||||
#ifdef _WIN32
|
||||
@@ -1553,8 +1564,13 @@ static void* http_worker(void* arg) {
|
||||
int fd = a->fd;
|
||||
#endif
|
||||
free(a);
|
||||
char *method = NULL, *path = NULL, *body = NULL;
|
||||
if (http_read_request(fd, &method, &path, &body, NULL) == 0) {
|
||||
char *method = NULL, *path = NULL, *body = NULL, *hdr_block = NULL;
|
||||
if (http_read_request(fd, &method, &path, &body, &hdr_block) == 0
|
||||
&& !el_http_request_authorized(method, path, hdr_block)) {
|
||||
/* Loopback hardening: EL_HTTP_AUTH_KEY is set and this request lacks the
|
||||
* matching X-Neuron-Auth header — refuse before it reaches any handler. */
|
||||
el_http_send_401(fd);
|
||||
} else if (method != NULL) {
|
||||
http_handler_fn h = http_lookup_active();
|
||||
char* response = NULL;
|
||||
/* HEAD: dispatch as GET so existing handlers respond with the same
|
||||
@@ -1582,7 +1598,7 @@ static void* http_worker(void* arg) {
|
||||
_tl_http_head_only = 0;
|
||||
free(response);
|
||||
}
|
||||
free(method); free(path); free(body);
|
||||
free(method); free(path); free(body); free(hdr_block);
|
||||
el_closesocket(fd);
|
||||
/* release a slot */
|
||||
pthread_mutex_lock(&_http_conn_mu);
|
||||
@@ -1592,6 +1608,108 @@ static void* http_worker(void* arg) {
|
||||
return NULL;
|
||||
}
|
||||
|
||||
/* ── loopback lock + local API-key auth (shipped desktop hardening) ────────
|
||||
* Both controls are OFF by default (their env vars unset), so dev, self-host,
|
||||
* and server builds behave exactly as before. The shipped macOS launcher
|
||||
* neuron-daemons.sh sets them so a customer's soul is neither reachable from
|
||||
* other machines on the LAN nor callable by other local users/processes
|
||||
* without the per-install key held in the login Keychain:
|
||||
*
|
||||
* EL_HTTP_BIND_HOST=127.0.0.1 -> bind loopback only (el_http_apply_bind_addr)
|
||||
* EL_HTTP_AUTH_KEY=<per-install> -> require "X-Neuron-Auth: <key>" per request
|
||||
*/
|
||||
|
||||
/* Set the listen address on the dual-stack (AF_INET6, V6ONLY=0) socket. Default
|
||||
* is in6addr_any (all interfaces) — unchanged. When EL_HTTP_BIND_HOST names a
|
||||
* loopback ("127.0.0.1", "localhost", "loopback", or "::1") we bind the IPv4-
|
||||
* mapped IPv6 loopback ::ffff:127.0.0.1: on a V6ONLY=0 socket this accepts IPv4
|
||||
* 127.0.0.1 clients (the desktop app connects there) while refusing every
|
||||
* off-machine address. */
|
||||
static void el_http_apply_bind_addr(struct sockaddr_in6* addr) {
|
||||
const char* h = getenv("EL_HTTP_BIND_HOST");
|
||||
int loopback = h && *h && (strcmp(h, "127.0.0.1") == 0
|
||||
|| strcmp(h, "localhost") == 0
|
||||
|| strcmp(h, "loopback") == 0
|
||||
|| strcmp(h, "::1") == 0);
|
||||
if (loopback) {
|
||||
memset(&addr->sin6_addr, 0, sizeof(addr->sin6_addr));
|
||||
addr->sin6_addr.s6_addr[10] = 0xff; /* ::ffff:127.0.0.1 */
|
||||
addr->sin6_addr.s6_addr[11] = 0xff;
|
||||
addr->sin6_addr.s6_addr[12] = 127;
|
||||
addr->sin6_addr.s6_addr[15] = 1;
|
||||
} else {
|
||||
addr->sin6_addr = in6addr_any;
|
||||
}
|
||||
}
|
||||
|
||||
/* Human-readable description of the active bind host, for the listen log line. */
|
||||
static const char* el_http_bind_desc(void) {
|
||||
const char* h = getenv("EL_HTTP_BIND_HOST");
|
||||
if (h && *h && (strcmp(h, "127.0.0.1") == 0 || strcmp(h, "localhost") == 0
|
||||
|| strcmp(h, "loopback") == 0 || strcmp(h, "::1") == 0)) {
|
||||
return "127.0.0.1 (loopback)";
|
||||
}
|
||||
return "[::] (dual-stack)";
|
||||
}
|
||||
|
||||
/* Case-insensitive compare of the first n bytes of a and b. */
|
||||
static int el_ci_eq_n(const char* a, const char* b, size_t n) {
|
||||
for (size_t i = 0; i < n; i++) {
|
||||
unsigned char ca = (unsigned char)a[i], cb = (unsigned char)b[i];
|
||||
if (tolower(ca) != tolower(cb)) return 0;
|
||||
}
|
||||
return 1;
|
||||
}
|
||||
|
||||
/* Return 1 iff the raw header block carries a header named `name` (case-
|
||||
* insensitive) whose trimmed value equals `want` exactly. */
|
||||
static int el_http_header_equals(const char* hdr_block, const char* name,
|
||||
const char* want) {
|
||||
if (!hdr_block || !name || !want) return 0;
|
||||
size_t nlen = strlen(name), wlen = strlen(want);
|
||||
const char* p = hdr_block;
|
||||
while (*p) {
|
||||
const char* line_end = strstr(p, "\r\n");
|
||||
const char* end = line_end ? line_end : p + strlen(p);
|
||||
const char* colon = memchr(p, ':', (size_t)(end - p));
|
||||
if (colon && (size_t)(colon - p) == nlen && el_ci_eq_n(p, name, nlen)) {
|
||||
const char* v = colon + 1;
|
||||
while (v < end && (*v == ' ' || *v == '\t')) v++;
|
||||
size_t vlen = (size_t)(end - v);
|
||||
while (vlen > 0 && (v[vlen - 1] == ' ' || v[vlen - 1] == '\t')) vlen--;
|
||||
if (vlen == wlen && memcmp(v, want, wlen) == 0) return 1;
|
||||
}
|
||||
if (!line_end) break;
|
||||
p = line_end + 2;
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
/* Authorize an inbound request. Enforcement is active only when EL_HTTP_AUTH_KEY
|
||||
* is set; otherwise every request is allowed (dev default). GET/HEAD /health*
|
||||
* are always allowed so launch-agent liveness probes work without the key. */
|
||||
static int el_http_request_authorized(const char* method, const char* path,
|
||||
const char* hdr_block) {
|
||||
const char* key = getenv("EL_HTTP_AUTH_KEY");
|
||||
if (!key || !*key) return 1;
|
||||
if (method && (strcmp(method, "GET") == 0 || strcmp(method, "HEAD") == 0)
|
||||
&& path && strncmp(path, "/health", 7) == 0) return 1;
|
||||
return el_http_header_equals(hdr_block, "x-neuron-auth", key);
|
||||
}
|
||||
|
||||
/* Minimal 401 for unauthorized requests — never reaches an EL handler. */
|
||||
static void el_http_send_401(int fd) {
|
||||
static const char* body = "{\"error\":\"unauthorized\",\"code\":\"auth_required\"}";
|
||||
char resp[256];
|
||||
int n = snprintf(resp, sizeof(resp),
|
||||
"HTTP/1.1 401 Unauthorized\r\n"
|
||||
"Content-Type: application/json\r\n"
|
||||
"Content-Length: %zu\r\n"
|
||||
"Connection: close\r\n\r\n%s",
|
||||
strlen(body), body);
|
||||
if (n > 0) http_send_all(fd, resp, (size_t)n);
|
||||
}
|
||||
|
||||
el_val_t http_serve(el_val_t port, el_val_t handler) {
|
||||
/* If `handler` looks like a string name, register it as the active handler. */
|
||||
const char* hname = EL_CSTR(handler);
|
||||
@@ -1610,13 +1728,13 @@ el_val_t http_serve(el_val_t port, el_val_t handler) {
|
||||
struct sockaddr_in6 addr;
|
||||
memset(&addr, 0, sizeof(addr));
|
||||
addr.sin6_family = AF_INET6;
|
||||
addr.sin6_addr = in6addr_any;
|
||||
el_http_apply_bind_addr(&addr);
|
||||
addr.sin6_port = htons((uint16_t)p);
|
||||
if (bind(sock, (struct sockaddr*)&addr, sizeof(addr)) < 0) {
|
||||
perror("bind"); el_closesocket(sock); return 0;
|
||||
}
|
||||
if (listen(sock, 64) < 0) { perror("listen"); el_closesocket(sock); return 0; }
|
||||
fprintf(stderr, "[http] listening on [::]:%d (dual-stack)\n", p);
|
||||
fprintf(stderr, "[http] listening on %s port %d\n", el_http_bind_desc(), p);
|
||||
while (1) {
|
||||
struct sockaddr_in6 cli;
|
||||
socklen_t clen = sizeof(cli);
|
||||
@@ -1866,13 +1984,13 @@ el_val_t http_serve_v2(el_val_t port, el_val_t handler) {
|
||||
struct sockaddr_in6 addr;
|
||||
memset(&addr, 0, sizeof(addr));
|
||||
addr.sin6_family = AF_INET6;
|
||||
addr.sin6_addr = in6addr_any;
|
||||
el_http_apply_bind_addr(&addr);
|
||||
addr.sin6_port = htons((uint16_t)p);
|
||||
if (bind(sock, (struct sockaddr*)&addr, sizeof(addr)) < 0) {
|
||||
perror("bind"); el_closesocket(sock); return 0;
|
||||
}
|
||||
if (listen(sock, 64) < 0) { perror("listen"); el_closesocket(sock); return 0; }
|
||||
fprintf(stderr, "[http v2] listening on [::]:%d (dual-stack)\n", p);
|
||||
fprintf(stderr, "[http v2] listening on %s port %d\n", el_http_bind_desc(), p);
|
||||
while (1) {
|
||||
struct sockaddr_in6 cli;
|
||||
socklen_t clen = sizeof(cli);
|
||||
@@ -1968,13 +2086,13 @@ void http_serve_async(el_val_t port, el_val_t handler) {
|
||||
struct sockaddr_in6 addr;
|
||||
memset(&addr, 0, sizeof(addr));
|
||||
addr.sin6_family = AF_INET6;
|
||||
addr.sin6_addr = in6addr_any;
|
||||
el_http_apply_bind_addr(&addr);
|
||||
addr.sin6_port = htons((uint16_t)p);
|
||||
if (bind(sock, (struct sockaddr*)&addr, sizeof(addr)) < 0) {
|
||||
perror("bind"); close(sock); return;
|
||||
}
|
||||
if (listen(sock, 64) < 0) { perror("listen"); close(sock); return; }
|
||||
fprintf(stderr, "[http] async listening on [::]:%d (dual-stack)\n", p);
|
||||
fprintf(stderr, "[http] async listening on %s port %d\n", el_http_bind_desc(), p);
|
||||
HttpServeAsyncArg* a = malloc(sizeof(HttpServeAsyncArg));
|
||||
if (!a) { close(sock); return; }
|
||||
a->sock = sock;
|
||||
@@ -3139,10 +3257,72 @@ static char* jp_parse_string_raw(JsonParser* jp) {
|
||||
case 'r': c = '\r'; break;
|
||||
case 't': c = '\t'; break;
|
||||
case 'u': {
|
||||
/* Skip 4 hex digits; emit '?' as a placeholder */
|
||||
for (int i = 0; i < 4 && jp->p < jp->end; i++) jp->p++;
|
||||
c = '?';
|
||||
break;
|
||||
/* Decode \uXXXX (with surrogate pairs) to UTF-8.
|
||||
* Ported from lang/releases/v1.0.0-20260501 (2026-08-08
|
||||
* self-review). This copy carried the identical defect:
|
||||
* the escape was skipped and a literal '?' emitted, which
|
||||
* silently destroyed every non-ASCII character in any JSON
|
||||
* string entering the runtime. Two copies of one parser
|
||||
* bug is exactly how this class of fault survives, so the
|
||||
* fix lands in both. See the release copy for the full
|
||||
* measurement and rationale. */
|
||||
unsigned cp = 0;
|
||||
int ok = 1;
|
||||
for (int i = 0; i < 4; i++) {
|
||||
if (jp->p >= jp->end) { ok = 0; break; }
|
||||
char h = *jp->p++;
|
||||
unsigned d;
|
||||
if (h >= '0' && h <= '9') d = (unsigned)(h - '0');
|
||||
else if (h >= 'a' && h <= 'f') d = (unsigned)(h - 'a' + 10);
|
||||
else if (h >= 'A' && h <= 'F') d = (unsigned)(h - 'A' + 10);
|
||||
else { ok = 0; break; }
|
||||
cp = (cp << 4) | d;
|
||||
}
|
||||
if (!ok) { c = '?'; break; }
|
||||
if (cp >= 0xD800 && cp <= 0xDBFF &&
|
||||
(size_t)(jp->end - jp->p) >= 6 &&
|
||||
jp->p[0] == '\\' && jp->p[1] == 'u') {
|
||||
const char* save = jp->p;
|
||||
unsigned lo = 0; int ok2 = 1;
|
||||
jp->p += 2;
|
||||
for (int i = 0; i < 4; i++) {
|
||||
char h = *jp->p++;
|
||||
unsigned d;
|
||||
if (h >= '0' && h <= '9') d = (unsigned)(h - '0');
|
||||
else if (h >= 'a' && h <= 'f') d = (unsigned)(h - 'a' + 10);
|
||||
else if (h >= 'A' && h <= 'F') d = (unsigned)(h - 'A' + 10);
|
||||
else { ok2 = 0; break; }
|
||||
lo = (lo << 4) | d;
|
||||
}
|
||||
if (ok2 && lo >= 0xDC00 && lo <= 0xDFFF)
|
||||
cp = 0x10000u + ((cp - 0xD800u) << 10) + (lo - 0xDC00u);
|
||||
else jp->p = save;
|
||||
}
|
||||
if (cp >= 0xD800 && cp <= 0xDFFF) cp = 0xFFFD;
|
||||
|
||||
char ub[4]; int un;
|
||||
if (cp < 0x80) {
|
||||
ub[0] = (char)cp; un = 1;
|
||||
} else if (cp < 0x800) {
|
||||
ub[0] = (char)(0xC0 | (cp >> 6));
|
||||
ub[1] = (char)(0x80 | (cp & 0x3F)); un = 2;
|
||||
} else if (cp < 0x10000) {
|
||||
ub[0] = (char)(0xE0 | (cp >> 12));
|
||||
ub[1] = (char)(0x80 | ((cp >> 6) & 0x3F));
|
||||
ub[2] = (char)(0x80 | (cp & 0x3F)); un = 3;
|
||||
} else {
|
||||
ub[0] = (char)(0xF0 | (cp >> 18));
|
||||
ub[1] = (char)(0x80 | ((cp >> 12) & 0x3F));
|
||||
ub[2] = (char)(0x80 | ((cp >> 6) & 0x3F));
|
||||
ub[3] = (char)(0x80 | (cp & 0x3F)); un = 4;
|
||||
}
|
||||
while (len + (size_t)un >= cap) {
|
||||
cap *= 2;
|
||||
out = realloc(out, cap);
|
||||
if (!out) { fputs("el_runtime: out of memory\n", stderr); exit(1); }
|
||||
}
|
||||
for (int i = 0; i < un; i++) out[len++] = ub[i];
|
||||
continue; /* bytes already appended */
|
||||
}
|
||||
default: c = esc; break;
|
||||
}
|
||||
@@ -6048,6 +6228,13 @@ void el_cgi_init(el_val_t name, el_val_t dharma_id, el_val_t principal,
|
||||
#define ENGRAM_SUPPRESSION_BREAKTHROUGH 5
|
||||
#define ENGRAM_BREAKTHROUGH_WEIGHT 0.25
|
||||
#define ENGRAM_INHIBITION_FACTOR 0.1
|
||||
/* ENGRAM_WM_CAP: hard global ceiling on nodes holding working_memory_weight
|
||||
* > 0 at any time. Cowan (2001) puts human WM capacity at ~4 chunks; 24 gives
|
||||
* the daemon generous headroom while preventing the unbounded growth observed
|
||||
* in production (wm_active 288-778 per heartbeat — "working memory" that is
|
||||
* really the whole recently-touched graph). Ported from release runtime
|
||||
* v1.0.0-20260501 Pass 5 on 2026-07-15 self-review. */
|
||||
#define ENGRAM_WM_CAP 24
|
||||
|
||||
/* ── Layered consciousness architecture ──────────────────────────────────────
|
||||
*
|
||||
@@ -6826,116 +7013,75 @@ static int istr_contains(const char* hay, const char* needle) {
|
||||
return 0;
|
||||
}
|
||||
|
||||
/* ---- SPEC-SEARCH-UPGRADE-OURS-2026-07-14: ranked search (BM25 + recency) ----
|
||||
* Replaces first-N-in-storage-order substring matching (measured 13% hit@5 on
|
||||
* the 15-query pinned eval; ranked model measured 93% offline). Deterministic,
|
||||
* local, transparent — no model call on the hot path. Multi-word queries score
|
||||
* per-token (rare+concentrated terms weigh most); ties break newest-first so
|
||||
* fresh memories stop losing to storage order. The transparent-layer identity
|
||||
* filter is preserved unchanged: hidden self layers stay invisible here and
|
||||
* surface only via engram_activate — the legitimate path. */
|
||||
/* ── Tokenized query matching ───────────────────────────────────────────
|
||||
* The engram query surface (search / activate / goal-bias) historically
|
||||
* matched the ENTIRE raw query string as a single case-insensitive
|
||||
* substring via istr_contains(field, q). That is Ctrl-F, not search:
|
||||
* a multi-word query like "windows msi signing" only matched a node whose
|
||||
* text contained that exact contiguous run, so real multi-word queries
|
||||
* returned zero. istr_contains stays as the per-TOKEN primitive; these
|
||||
* helpers split the query on whitespace and match ANY token, then rank by
|
||||
* how many DISTINCT tokens a node covers. Single-token queries are a strict
|
||||
* special case (score is 0 or 1) so single-word callers never regress. */
|
||||
#define ENGRAM_MAX_QTOKENS 32
|
||||
#define ENGRAM_QTOK_LEN 256
|
||||
|
||||
#define ENGRAM_BM25_MAX_QTOK 16
|
||||
#define ENGRAM_BM25_TOKLEN 48
|
||||
|
||||
static int engram_tok_next(const char** ps, char* out, int cap) {
|
||||
const char* s = *ps;
|
||||
while (*s && !isalnum((unsigned char)*s)) s++;
|
||||
if (!*s) { *ps = s; return 0; }
|
||||
/* Split q on whitespace into up to ENGRAM_MAX_QTOKENS distinct
|
||||
* (case-insensitive) tokens. Returns the token count. Over-long tokens are
|
||||
* truncated to ENGRAM_QTOK_LEN-1; over-count tokens are ignored. */
|
||||
static int engram_tokenize_query(const char* q,
|
||||
char toks[][ENGRAM_QTOK_LEN], int maxtok) {
|
||||
int n = 0;
|
||||
while (*s && isalnum((unsigned char)*s)) {
|
||||
if (n < cap - 1) out[n++] = (char)tolower((unsigned char)*s);
|
||||
s++;
|
||||
if (!q) return 0;
|
||||
const char* p = q;
|
||||
while (*p && n < maxtok) {
|
||||
while (*p && isspace((unsigned char)*p)) p++;
|
||||
if (!*p) break;
|
||||
char buf[ENGRAM_QTOK_LEN];
|
||||
size_t tl = 0;
|
||||
while (*p && !isspace((unsigned char)*p)) {
|
||||
if (tl < sizeof(buf) - 1) buf[tl++] = *p;
|
||||
p++;
|
||||
}
|
||||
buf[tl] = '\0';
|
||||
if (tl == 0) continue;
|
||||
int dup = 0;
|
||||
for (int s = 0; s < n; s++) {
|
||||
if (strcasecmp(toks[s], buf) == 0) { dup = 1; break; }
|
||||
}
|
||||
if (dup) continue;
|
||||
memcpy(toks[n], buf, tl + 1);
|
||||
n++;
|
||||
}
|
||||
out[n] = 0; *ps = s; return 1;
|
||||
return n;
|
||||
}
|
||||
|
||||
static void engram_field_stats(const char* field,
|
||||
char qtok[][ENGRAM_BM25_TOKLEN], int nq,
|
||||
int64_t* tf, int64_t* doclen) {
|
||||
if (!field) return;
|
||||
char buf[ENGRAM_BM25_TOKLEN];
|
||||
const char* p = field;
|
||||
while (engram_tok_next(&p, buf, sizeof buf)) {
|
||||
(*doclen)++;
|
||||
for (int t = 0; t < nq; t++)
|
||||
if (strcmp(buf, qtok[t]) == 0) tf[t]++;
|
||||
/* Count how many of the ntok distinct query tokens appear (case-insensitive)
|
||||
* in the node's content, label, or tags. 0 == no match. */
|
||||
static int engram_node_match_score(const EngramNode* n,
|
||||
char toks[][ENGRAM_QTOK_LEN], int ntok) {
|
||||
int score = 0;
|
||||
for (int t = 0; t < ntok; t++) {
|
||||
if (istr_contains(n->content, toks[t]) ||
|
||||
istr_contains(n->label, toks[t]) ||
|
||||
istr_contains(n->tags, toks[t]))
|
||||
score++;
|
||||
}
|
||||
return score;
|
||||
}
|
||||
|
||||
typedef struct { double score; int64_t created; int64_t idx; } EngramHit;
|
||||
|
||||
static int engram_hit_cmp(const void* a, const void* b) {
|
||||
const EngramHit* x = (const EngramHit*)a;
|
||||
const EngramHit* y = (const EngramHit*)b;
|
||||
if (x->score != y->score) return (x->score < y->score) ? 1 : -1;
|
||||
if (x->created != y->created) return (x->created < y->created) ? 1 : -1;
|
||||
/* Rank entry: distinct-token match count (primary, desc) then salience
|
||||
* (tiebreak, desc). */
|
||||
typedef struct { int64_t idx; int score; double salience; } EngramRankEntry;
|
||||
static int engram_rank_cmp(const void* a, const void* b) {
|
||||
const EngramRankEntry* ea = (const EngramRankEntry*)a;
|
||||
const EngramRankEntry* eb = (const EngramRankEntry*)b;
|
||||
if (ea->score != eb->score) return eb->score - ea->score; /* desc */
|
||||
if (ea->salience < eb->salience) return 1;
|
||||
if (ea->salience > eb->salience) return -1;
|
||||
return 0;
|
||||
}
|
||||
|
||||
/* Scores every visible node against the query; writes ranked hits into `out`
|
||||
* (caller allocates g->node_count entries). Returns min(hits, lim). */
|
||||
static int64_t engram_search_ranked(EngramStore* g, const char* q, int64_t lim,
|
||||
EngramHit* out) {
|
||||
char qtok[ENGRAM_BM25_MAX_QTOK][ENGRAM_BM25_TOKLEN];
|
||||
int nq = 0;
|
||||
{
|
||||
const char* p = q; char buf[ENGRAM_BM25_TOKLEN];
|
||||
while (nq < ENGRAM_BM25_MAX_QTOK && engram_tok_next(&p, buf, sizeof buf)) {
|
||||
int dup = 0;
|
||||
for (int t = 0; t < nq; t++)
|
||||
if (strcmp(qtok[t], buf) == 0) { dup = 1; break; }
|
||||
if (!dup) { strcpy(qtok[nq], buf); nq++; }
|
||||
}
|
||||
}
|
||||
if (nq == 0) return 0;
|
||||
|
||||
int64_t N = g->node_count;
|
||||
int64_t* tfm = (int64_t*)calloc((size_t)(N * nq), sizeof(int64_t));
|
||||
int64_t* dlen = (int64_t*)calloc((size_t)N, sizeof(int64_t));
|
||||
if (!tfm || !dlen) { free(tfm); free(dlen); return 0; }
|
||||
int64_t df[ENGRAM_BM25_MAX_QTOK] = {0};
|
||||
double total_len = 0.0; int64_t live = 0;
|
||||
for (int64_t i = 0; i < N; i++) {
|
||||
EngramNode* n = &g->nodes[i];
|
||||
if (engram_layer_is_transparent(n->layer_id)) continue;
|
||||
live++;
|
||||
int64_t* tf = &tfm[i * nq];
|
||||
engram_field_stats(n->content, qtok, nq, tf, &dlen[i]);
|
||||
engram_field_stats(n->label, qtok, nq, tf, &dlen[i]);
|
||||
engram_field_stats(n->tags, qtok, nq, tf, &dlen[i]);
|
||||
total_len += (double)dlen[i];
|
||||
for (int t = 0; t < nq; t++) if (tf[t] > 0) df[t]++;
|
||||
}
|
||||
double avg = (live > 0) ? total_len / (double)live : 1.0;
|
||||
if (avg <= 0.0) avg = 1.0;
|
||||
const double k1 = 1.2, b = 0.75;
|
||||
int64_t nhits = 0;
|
||||
for (int64_t i = 0; i < N; i++) {
|
||||
EngramNode* n = &g->nodes[i];
|
||||
if (engram_layer_is_transparent(n->layer_id)) continue;
|
||||
int64_t* tf = &tfm[i * nq];
|
||||
double s = 0.0;
|
||||
for (int t = 0; t < nq; t++) {
|
||||
if (tf[t] == 0) continue;
|
||||
double idf = log(((double)live - (double)df[t] + 0.5) /
|
||||
((double)df[t] + 0.5) + 1.0);
|
||||
double tfd = (double)tf[t];
|
||||
s += idf * (tfd * (k1 + 1.0)) /
|
||||
(tfd + k1 * (1.0 - b + b * (double)dlen[i] / avg));
|
||||
}
|
||||
if (s > 0.0) {
|
||||
out[nhits].score = s;
|
||||
out[nhits].created = n->created_at;
|
||||
out[nhits].idx = i;
|
||||
nhits++;
|
||||
}
|
||||
}
|
||||
free(tfm); free(dlen);
|
||||
qsort(out, (size_t)nhits, sizeof(EngramHit), engram_hit_cmp);
|
||||
return (nhits < lim) ? nhits : lim;
|
||||
}
|
||||
|
||||
el_val_t engram_search(el_val_t query, el_val_t limit) {
|
||||
EngramStore* g = engram_get();
|
||||
const char* q = EL_CSTR(query);
|
||||
@@ -6943,12 +7089,33 @@ el_val_t engram_search(el_val_t query, el_val_t limit) {
|
||||
if (lim <= 0) lim = 100;
|
||||
el_val_t lst = el_list_empty();
|
||||
if (!q || !*q) return lst;
|
||||
if (g->node_count == 0) return lst;
|
||||
EngramHit* hits = (EngramHit*)malloc((size_t)g->node_count * sizeof(EngramHit));
|
||||
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
|
||||
int ntok = engram_tokenize_query(q, toks, ENGRAM_MAX_QTOKENS);
|
||||
if (ntok == 0) return lst;
|
||||
EngramRankEntry* hits = malloc((size_t)g->node_count * sizeof(EngramRankEntry));
|
||||
if (!hits) return lst;
|
||||
int64_t k = engram_search_ranked(g, q, lim, hits);
|
||||
for (int64_t i = 0; i < k; i++)
|
||||
lst = el_list_append(lst, engram_node_to_map(&g->nodes[hits[i].idx]));
|
||||
int64_t nhits = 0;
|
||||
for (int64_t i = 0; i < g->node_count; i++) {
|
||||
EngramNode* n = &g->nodes[i];
|
||||
/* Filter transparent layers: nodes whose layer is `transparent=1`
|
||||
* shape output but are invisible to introspection ("what do you
|
||||
* know about yourself"). They still surface via engram_activate
|
||||
* + engram_compile_layered_json — that's the legitimate path. */
|
||||
if (engram_layer_is_transparent(n->layer_id)) continue;
|
||||
int sc = engram_node_match_score(n, toks, ntok);
|
||||
if (sc > 0) {
|
||||
hits[nhits].idx = i;
|
||||
hits[nhits].score = sc;
|
||||
hits[nhits].salience = n->salience;
|
||||
nhits++;
|
||||
}
|
||||
}
|
||||
/* Rank by distinct tokens matched (desc) then salience (desc), then cap. */
|
||||
qsort(hits, (size_t)nhits, sizeof(EngramRankEntry), engram_rank_cmp);
|
||||
int64_t end = nhits < lim ? nhits : lim;
|
||||
for (int64_t k = 0; k < end; k++) {
|
||||
lst = el_list_append(lst, engram_node_to_map(&g->nodes[hits[k].idx]));
|
||||
}
|
||||
free(hits);
|
||||
return lst;
|
||||
}
|
||||
@@ -7226,10 +7393,14 @@ static double engram_temporal_proximity_bonus(int64_t node_created,
|
||||
static double engram_goal_bias(const EngramNode* n, const char* query) {
|
||||
if (!query || !*query) return 1.0;
|
||||
double bias = 1.0;
|
||||
/* Direct lexical overlap: node content/label/tags share text with query. */
|
||||
if (istr_contains(n->content, query) || istr_contains(n->label, query) ||
|
||||
istr_contains(n->tags, query)) {
|
||||
bias += 0.5;
|
||||
/* Direct lexical overlap, graded by token coverage: a node covering all
|
||||
* query tokens gets the full +0.5; partial coverage gets a proportional
|
||||
* share. Single-token queries → full +0.5 on match, identical to before. */
|
||||
{
|
||||
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
|
||||
int ntok = engram_tokenize_query(query, toks, ENGRAM_MAX_QTOKENS);
|
||||
int sc = engram_node_match_score(n, toks, ntok);
|
||||
if (sc > 0 && ntok > 0) bias += 0.5 * ((double)sc / (double)ntok);
|
||||
}
|
||||
/* Node-type resonance with query intent. */
|
||||
int technical_query = istr_contains(query, "code") ||
|
||||
@@ -7268,6 +7439,48 @@ static double engram_goal_bias(const EngramNode* n, const char* query) {
|
||||
return bias;
|
||||
}
|
||||
|
||||
/* eg_cmp_double_desc — qsort comparator, descending doubles. */
|
||||
static int eg_cmp_double_desc(const void* a, const void* b) {
|
||||
double da = *(const double*)a, db = *(const double*)b;
|
||||
if (da < db) return 1;
|
||||
if (da > db) return -1;
|
||||
return 0;
|
||||
}
|
||||
|
||||
/* eg_enforce_wm_cap_global — clamp the store-wide working-memory population
|
||||
* to ENGRAM_WM_CAP, keeping the top-K by current weight. Runs at every point
|
||||
* that materializes WM: post-activation persist and snapshot load/merge.
|
||||
* (Ported from release runtime v1.0.0-20260501 Pass 5, 2026-07-15.) */
|
||||
static void eg_enforce_wm_cap_global(EngramStore* g) {
|
||||
int64_t wm_count = 0;
|
||||
for (int64_t i = 0; i < g->node_count; i++) {
|
||||
if (g->nodes[i].working_memory_weight > 0.0) wm_count++;
|
||||
}
|
||||
if (wm_count <= ENGRAM_WM_CAP) return;
|
||||
double* vals = malloc((size_t)wm_count * sizeof(double));
|
||||
if (!vals) return; /* OOM: over cap this call, no corruption */
|
||||
int64_t vi = 0;
|
||||
for (int64_t i = 0; i < g->node_count; i++) {
|
||||
if (g->nodes[i].working_memory_weight > 0.0)
|
||||
vals[vi++] = g->nodes[i].working_memory_weight;
|
||||
}
|
||||
qsort(vals, (size_t)wm_count, sizeof(double), eg_cmp_double_desc);
|
||||
double cutoff = vals[ENGRAM_WM_CAP - 1];
|
||||
free(vals);
|
||||
int64_t above = 0;
|
||||
for (int64_t i = 0; i < g->node_count; i++) {
|
||||
if (g->nodes[i].working_memory_weight > cutoff) above++;
|
||||
}
|
||||
int64_t slots_at_cutoff = ENGRAM_WM_CAP - above;
|
||||
for (int64_t i = 0; i < g->node_count; i++) {
|
||||
EngramNode* n = &g->nodes[i];
|
||||
if (n->working_memory_weight <= 0.0) continue;
|
||||
if (n->working_memory_weight > cutoff) continue;
|
||||
if (slots_at_cutoff > 0) { slots_at_cutoff--; continue; }
|
||||
n->working_memory_weight = 0.0; /* evict: over global cap */
|
||||
}
|
||||
}
|
||||
|
||||
el_val_t engram_activate(el_val_t query, el_val_t depth) {
|
||||
EngramStore* g = engram_get();
|
||||
const char* q = EL_CSTR(query);
|
||||
@@ -7295,14 +7508,21 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
|
||||
if (!seeds) {
|
||||
free(best_bg); free(best_hops); free(reached); return out;
|
||||
}
|
||||
/* Tokenize once: a node seeds if it matches ANY query token, and its seed
|
||||
* activation is scaled by token coverage (fraction of distinct query
|
||||
* tokens it contains) so a node matching all words seeds more strongly
|
||||
* than one matching a single word. Single-word queries → coverage 1.0,
|
||||
* identical to the prior whole-query behavior. */
|
||||
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
|
||||
int ntok = engram_tokenize_query(q, toks, ENGRAM_MAX_QTOKENS);
|
||||
for (int64_t i = 0; i < g->node_count; i++) {
|
||||
EngramNode* n = &g->nodes[i];
|
||||
if (istr_contains(n->content, q) ||
|
||||
istr_contains(n->label, q) ||
|
||||
istr_contains(n->tags, q)) {
|
||||
int sc = engram_node_match_score(n, toks, ntok);
|
||||
if (sc > 0) {
|
||||
double tdecay = engram_temporal_decay(n, now_ms);
|
||||
double dampen = engram_activation_dampen(n);
|
||||
double act = n->salience * tdecay * dampen;
|
||||
double cover = ntok > 0 ? (double)sc / (double)ntok : 1.0;
|
||||
double act = n->salience * tdecay * dampen * cover;
|
||||
seeds[seed_count].idx = i;
|
||||
seeds[seed_count].act = act;
|
||||
seeds[seed_count].created_at = n->created_at;
|
||||
@@ -7466,6 +7686,12 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
|
||||
g->nodes[i].working_memory_weight = wm_weights[i];
|
||||
}
|
||||
|
||||
/* Global WM cap: keep only the top ENGRAM_WM_CAP by weight across the
|
||||
* whole store (see eg_enforce_wm_cap_global). Without this, repeated
|
||||
* activation calls accumulate hundreds of "promoted" nodes and WM stops
|
||||
* meaning anything (production heartbeats showed wm_active up to 778). */
|
||||
eg_enforce_wm_cap_global(g);
|
||||
|
||||
/* ── Collect all background-activated nodes for the return value ────
|
||||
* Callers see both layers. Context compilation uses only promoted nodes
|
||||
* (working_memory_weight > 0). Sort: promoted first by wm_weight desc,
|
||||
@@ -7843,6 +8069,9 @@ el_val_t engram_load(el_val_t path) {
|
||||
}
|
||||
}
|
||||
}
|
||||
/* WM cap discipline applies to every entry point that materializes WM,
|
||||
* including snapshot restore (see eg_enforce_wm_cap_global). */
|
||||
eg_enforce_wm_cap_global(g);
|
||||
free(data);
|
||||
return 1;
|
||||
}
|
||||
@@ -7863,24 +8092,72 @@ el_val_t engram_get_node_json(el_val_t id) {
|
||||
return el_wrap_str(jb_finish(&b));
|
||||
}
|
||||
|
||||
/* engram_get_node_by_label — find the first node whose label field exactly
|
||||
* matches the given string. Returns the node as a JSON object string, or "{}"
|
||||
* if no match is found.
|
||||
*
|
||||
* Exact match (strcmp, not substring) because labels like "conv:history"
|
||||
* must not collide with nodes whose content contains that substring.
|
||||
*
|
||||
* Ported from the release runtime 2026-07-16 self-review: chat.el has called
|
||||
* this since 2026-07-01 but the function only existed in
|
||||
* releases/v1.0.0-20260501/el_runtime.c — the soul daemon (which builds
|
||||
* against THIS runtime) failed to compile once clang made implicit
|
||||
* declarations an error. */
|
||||
el_val_t engram_get_node_by_label(el_val_t label) {
|
||||
const char* lbl = EL_CSTR(label);
|
||||
if (!lbl || !*lbl) return el_wrap_str(el_strdup("{}"));
|
||||
EngramStore* g = engram_get();
|
||||
for (int64_t i = 0; i < g->node_count; i++) {
|
||||
EngramNode* n = &g->nodes[i];
|
||||
if (n->label && strcmp(n->label, lbl) == 0) {
|
||||
JsonBuf b; jb_init(&b);
|
||||
engram_emit_node_json(&b, n);
|
||||
return el_wrap_str(jb_finish(&b));
|
||||
}
|
||||
}
|
||||
return el_wrap_str(el_strdup("{}"));
|
||||
}
|
||||
|
||||
el_val_t engram_search_json(el_val_t query, el_val_t limit) {
|
||||
/* SPEC-SEARCH-UPGRADE 2026-07-14: same ranked BM25+recency core as
|
||||
* engram_search; transparent-layer identity filter enforced inside it. */
|
||||
EngramStore* g = engram_get();
|
||||
const char* q = EL_CSTR(query);
|
||||
int64_t lim = (int64_t)limit;
|
||||
if (lim <= 0) lim = 100;
|
||||
JsonBuf b; jb_init(&b);
|
||||
jb_putc(&b, '[');
|
||||
if (q && *q && g->node_count > 0) {
|
||||
EngramHit* hits = (EngramHit*)malloc((size_t)g->node_count * sizeof(EngramHit));
|
||||
if (hits) {
|
||||
int64_t k = engram_search_ranked(g, q, lim, hits);
|
||||
for (int64_t i = 0; i < k; i++) {
|
||||
if (i) jb_putc(&b, ',');
|
||||
engram_emit_node_json(&b, &g->nodes[hits[i].idx]);
|
||||
int first = 1;
|
||||
if (q && *q) {
|
||||
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
|
||||
int ntok = engram_tokenize_query(q, toks, ENGRAM_MAX_QTOKENS);
|
||||
if (ntok > 0) {
|
||||
EngramRankEntry* hits =
|
||||
malloc((size_t)g->node_count * sizeof(EngramRankEntry));
|
||||
if (hits) {
|
||||
int64_t nhits = 0;
|
||||
for (int64_t i = 0; i < g->node_count; i++) {
|
||||
EngramNode* n = &g->nodes[i];
|
||||
/* Filter transparent layers — same as engram_search. */
|
||||
if (engram_layer_is_transparent(n->layer_id)) continue;
|
||||
int sc = engram_node_match_score(n, toks, ntok);
|
||||
if (sc > 0) {
|
||||
hits[nhits].idx = i;
|
||||
hits[nhits].score = sc;
|
||||
hits[nhits].salience = n->salience;
|
||||
nhits++;
|
||||
}
|
||||
}
|
||||
/* Rank by distinct tokens matched (desc) then salience (desc). */
|
||||
qsort(hits, (size_t)nhits, sizeof(EngramRankEntry),
|
||||
engram_rank_cmp);
|
||||
int64_t end = nhits < lim ? nhits : lim;
|
||||
for (int64_t k = 0; k < end; k++) {
|
||||
if (!first) jb_putc(&b, ',');
|
||||
engram_emit_node_json(&b, &g->nodes[hits[k].idx]);
|
||||
first = 0;
|
||||
}
|
||||
free(hits);
|
||||
}
|
||||
free(hits);
|
||||
}
|
||||
}
|
||||
jb_putc(&b, ']');
|
||||
@@ -8399,6 +8676,9 @@ el_val_t engram_load_merge(el_val_t path) {
|
||||
}
|
||||
}
|
||||
|
||||
/* Merged nodes can carry snapshot WM weights too — hold the cap here as
|
||||
* well (see eg_enforce_wm_cap_global). */
|
||||
eg_enforce_wm_cap_global(g);
|
||||
free(data);
|
||||
return (el_val_t)added_nodes;
|
||||
}
|
||||
|
||||
@@ -632,6 +632,7 @@ el_val_t engram_load(el_val_t path);
|
||||
* can pass results straight through without round-tripping ElList/ElMap
|
||||
* through json_stringify. */
|
||||
el_val_t engram_get_node_json(el_val_t id);
|
||||
el_val_t engram_get_node_by_label(el_val_t label);
|
||||
el_val_t engram_search_json(el_val_t query, el_val_t limit);
|
||||
el_val_t engram_scan_nodes_json(el_val_t limit, el_val_t offset);
|
||||
el_val_t engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, el_val_t offset);
|
||||
|
||||
@@ -23,10 +23,29 @@ fn tok_at(tokens: [Any], pos: Int) -> Map<String, Any> {
|
||||
}
|
||||
|
||||
fn tok_kind(tokens: [Any], pos: Int) -> String {
|
||||
// Out-of-range reads must report the Eof sentinel so every `== "Eof"`
|
||||
// termination guard in the parser fires. Without this, reading past the
|
||||
// single trailing Eof token returns runtime null (el_list_get OOB -> 0),
|
||||
// which matches no delimiter, letting inner parse loops append AST nodes
|
||||
// forever on malformed input -> unbounded allocation -> OOM.
|
||||
let n: Int = native_list_len(tokens) / 2
|
||||
if pos < 0 {
|
||||
return "Eof"
|
||||
}
|
||||
if pos >= n {
|
||||
return "Eof"
|
||||
}
|
||||
native_list_get(tokens, pos * 2)
|
||||
}
|
||||
|
||||
fn tok_value(tokens: [Any], pos: Int) -> String {
|
||||
let n: Int = native_list_len(tokens) / 2
|
||||
if pos < 0 {
|
||||
return ""
|
||||
}
|
||||
if pos >= n {
|
||||
return ""
|
||||
}
|
||||
native_list_get(tokens, pos * 2 + 1)
|
||||
}
|
||||
|
||||
@@ -35,7 +54,12 @@ fn expect(tokens: [Any], pos: Int, kind: String) -> Int {
|
||||
if k == kind {
|
||||
return pos + 1
|
||||
}
|
||||
// On mismatch just advance; error recovery is best-effort
|
||||
// On mismatch, error recovery is best-effort. But never step PAST the Eof
|
||||
// sentinel: once at Eof a mismatch means the input ended early, and
|
||||
// advancing would run the cursor off the token list.
|
||||
if k == "Eof" {
|
||||
return pos
|
||||
}
|
||||
pos + 1
|
||||
}
|
||||
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -117,6 +117,15 @@ el_val_t el_min(el_val_t a, el_val_t b);
|
||||
void el_retain(el_val_t v);
|
||||
void el_release(el_val_t v);
|
||||
|
||||
/* ── Arena scoping ────────────────────────────────────────────────────────────
|
||||
* el_arena_push() activates the string arena (if not already active) and
|
||||
* returns a mark; el_arena_pop(mark) frees all strings allocated since that
|
||||
* mark. Used by codegen for per-function/statement scoping and by long-running
|
||||
* EL loops (e.g. the soul daemon's awareness tick) to reclaim per-iteration
|
||||
* allocations. */
|
||||
el_val_t el_arena_push(void);
|
||||
el_val_t el_arena_pop(el_val_t mark);
|
||||
|
||||
/* ── List ────────────────────────────────────────────────────────────────── */
|
||||
|
||||
el_val_t el_list_new(el_val_t count, ...);
|
||||
@@ -142,6 +151,7 @@ el_val_t http_get_with_headers(el_val_t url, el_val_t headers_map);
|
||||
el_val_t http_post_with_headers(el_val_t url, el_val_t body, el_val_t headers_map);
|
||||
el_val_t http_post_form_auth(el_val_t url, el_val_t form_body, el_val_t auth_header);
|
||||
el_val_t http_delete(el_val_t url);
|
||||
el_val_t http_delete_json(el_val_t url, el_val_t json_body);
|
||||
void http_serve(el_val_t port, el_val_t handler);
|
||||
void http_set_handler(el_val_t name);
|
||||
|
||||
@@ -167,6 +177,11 @@ void http_set_handler(el_val_t name);
|
||||
void http_serve_v2(el_val_t port, el_val_t handler);
|
||||
void http_set_handler_v2(el_val_t name);
|
||||
|
||||
/* Non-blocking variant of http_serve: runs the accept loop in a background
|
||||
* pthread and returns immediately so the caller can continue (used by the
|
||||
* soul daemon to run awareness_run() after starting its HTTP API). */
|
||||
void http_serve_async(el_val_t port, el_val_t handler);
|
||||
|
||||
/* Build an HTTP response envelope. `headers_json` should be a JSON object
|
||||
* literal like `{"WWW-Authenticate":"Basic"}` (or "" / "{}" for none). The
|
||||
* returned string carries the discriminator `{"el_http_response":1,...}`
|
||||
@@ -576,6 +591,7 @@ el_val_t engram_list_layers(void);
|
||||
el_val_t engram_get_node(el_val_t id);
|
||||
void engram_strengthen(el_val_t node_id);
|
||||
void engram_forget(el_val_t node_id);
|
||||
el_val_t engram_prune_telemetry(el_val_t older_than_ms);
|
||||
el_val_t engram_node_count(void);
|
||||
el_val_t engram_search(el_val_t query, el_val_t limit);
|
||||
el_val_t engram_scan_nodes(el_val_t limit, el_val_t offset);
|
||||
@@ -594,12 +610,32 @@ el_val_t engram_load(el_val_t path);
|
||||
* can pass results straight through without round-tripping ElList/ElMap
|
||||
* through json_stringify. */
|
||||
el_val_t engram_get_node_json(el_val_t id);
|
||||
el_val_t engram_get_node_by_label(el_val_t label);
|
||||
el_val_t engram_search_json(el_val_t query, el_val_t limit);
|
||||
el_val_t engram_scan_nodes_json(el_val_t limit, el_val_t offset);
|
||||
el_val_t engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, el_val_t offset);
|
||||
el_val_t engram_neighbors_json(el_val_t node_id, el_val_t max_depth, el_val_t direction);
|
||||
el_val_t engram_activate_json(el_val_t query, el_val_t depth);
|
||||
el_val_t engram_stats_json(void);
|
||||
el_val_t engram_act_stats_json(void);
|
||||
el_val_t engram_text_health_json(void);
|
||||
el_val_t engram_cosine_sim(el_val_t id_a, el_val_t id_b);
|
||||
/* Destructively pop up to `max` newly-formed Hebbian associations as a JSON
|
||||
* array of {from_id,to_id,weight,hebb}. The learning process (soul daemon) is
|
||||
* not the process that owns persistence (engram HTTP server); this is how a
|
||||
* self-formed association crosses that boundary. (2026-08-07 self-review.) */
|
||||
el_val_t engram_hebb_drain_json(el_val_t max);
|
||||
/* Document frequency of a term across node labels — term-specificity signal
|
||||
* for curiosity seed selection. (2026-08-03 self-review.) */
|
||||
el_val_t engram_label_df(el_val_t term);
|
||||
/* Best curiosity seed from one node: argmax over idf·position·casing across
|
||||
* the candidate tokens of its label, falling back to its content when the
|
||||
* label is a sentinel. Excludes pipe-delimited tabu terms during selection
|
||||
* and gates candidates to the df band [min_df, max_df]. Returns "" when
|
||||
* nothing qualifies. (2026-08-13 self-review.) */
|
||||
el_val_t engram_salient_term(el_val_t node_id, el_val_t max_df,
|
||||
el_val_t min_df, el_val_t tabu);
|
||||
el_val_t engram_embed_backfill(el_val_t count);
|
||||
el_val_t engram_list_layers_json(void);
|
||||
/* Working memory introspection — count, mean weight, and top-N snapshot.
|
||||
* Ported from el-compiler/runtime on 2026-06-30 self-review. */
|
||||
|
||||
Reference in New Issue
Block a user