Compare commits

..

3 Commits

Author SHA1 Message Date
Tim Lingo f34270d63d runtime: fs_read length hint must be paired with its buffer — fixes truncated HTTP responses
El SDK Release / build-and-release (pull_request) Failing after 25s
The binary-safe fs_read length (_tl_fs_read_len) was consumed by the HTTP
response path for ANY body, even when the handler wrapped the file into a
larger reply. Content-Length then lied AND the send stopped short: the
safety-contact routes returned 178 of 208/218 bytes, cut mid-'set_at' —
unparseable JSON. The desktop app read that as failure: fresh installs
trapped at 'Set your safety contact' (POST reply mangled) and configured
users saw the gate re-appear every launch (GET reply mangled). Worse, a
stale hint LARGER than a later body would over-read heap memory out the
socket.

Fix: pair the hint with the exact buffer pointer it describes; consume it
only when the response IS that buffer (binary file serving keeps working,
the hint follows the worker's copy); reset both at request start. Also
ports engram_get_node_by_label (from releases/v1.0.0) needed by soul.el
session continuity in local mode — not-found returns "" (matches shipped
behavior; '{}' flips the truthiness check upstream).

Verified: genesis boot + byte-math E2E on :7797 sandbox — safety-contact
GET/POST/GET all Content-Length==body, json-parse clean; /health,
/api/config regressions match.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 18:25:21 -05:00
Tim Lingo f76ccc0590 engram: ranked BM25+recency search replaces storage-order substring; URL-decode GET query params
Measured on the live container mind (pinned 40-query eval, judged): substring
2/40=5% hit@5 -> ranked 35/40=88%. Multi-word queries stop returning zero; new
memories stop losing to storage order (created_at tiebreak). Transparent-layer
identity filter preserved in both passes; jb_finish (#64) tail preserved.
query_param now url_decode()s values - %XX arrived literal before (pre-existing
GET defect, masked while multi-word substring returned nothing anyway).
E2E-verified in Tim's container deployment 2026-07-14/15; eval harness:
docs repo research-archive/p0-prototypes/eval_pinned_40q_20260715.py.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 07:48:09 -05:00
will.anderson 2b2a1246e7 Merge pull request 'runtime: fix the memory-leak + write-corruption pair in el_runtime.c' (#64) from hotfix/el-runtime-leak-and-persist into main
El SDK Release / build-and-release (push) Failing after 10m52s
2026-07-13 21:23:31 +00:00
51 changed files with 507 additions and 826562 deletions
-65
View File
@@ -1,65 +0,0 @@
# ELP language consolidation — full-lexicon backfill (stage)
Branch: `stage-elp-lang-consolidation` (stage-bound; NOT the live soul :8742).
Consolidates scattered Python language-realizer work (`~/Desktop/lang-realizers`,
`~/Desktop/lang-poetry-experiment`, `~/semitic_engine`) into the ELP `.el`
structure, generating **full lexicons** (complete UniMorph + kaikki.org
Wiktionary — real gender, real inflections) instead of the demo/curated subsets
the prototypes shipped.
## ELP before this branch
- 18 classical/ancient languages fully done (vocab + morphology + tests):
akk ang cop egy enm fro gez goh got grc non peo pi sa sga sux txb uga.
- 11 modern/classical languages had `morphology-<code>.el` in the build manifest
but **no vocabulary and no lang_profile**: es fr de ja ar he hi ru fi sw la.
- The ES port (`stage-elp-es-port`) had a *demo-scale* vocabulary-es.el (~350
entries, s-expr form).
## Landed on this branch (full-lexicon seed-fn format, matching the 18 ancients)
Vocabulary schema per row: `[lemma, pos, form0, form1, form2, en_gloss, hint]`.
Files are ELP runtime **seed data** (loaded via the Engram at runtime), so — like
all 18 classical `vocabulary-*.el` — they are intentionally NOT in the build
manifest. Syntax validated: the chunked `fn vocab_<code>_seed_pN` format
compiles cleanly to C via `elc` (correct UTF-8).
| code | in-ELP-morph? | vocab entries | verbs | nouns | adjs | profile |
|------|---------------|--------------:|------:|------:|-----:|---------|
| es | yes | 72,032 | 6,695 | 48,353 | 16,984 | yes |
| fr | yes | 130,517 | 7,534 | 77,344 | 45,639 | yes |
| de | yes | 144,692 | 6,661 | 133,162 | 4,869 | yes |
| la | yes | 22,590 | 82 | 13,436 | 9,072 | yes |
| it | no (bonus) | 193,675 | 10,008 | 109,459 | 74,208 | yes |
| pt | no (bonus) | 115,772 | 4,001 | 72,073 | 39,698 | yes |
| ro | no (bonus) | 86,504 | 1,216 | 65,915 | 19,373 | yes |
| ca | no (bonus) | 47,112 | 1,547 | 28,830 | 16,735 | yes |
|**total**| |**812,894** | | | | |
Generators (reproducible): `elp/tests/lang-gen/gen_elp_seed_full.py` (Romance),
`gen_elp_seed_de_la.py` (German declension + Latin case-paradigm mapping). They
read the pre-built morph caches in `~/Desktop/lang-realizers/data/` (UniMorph +
kaikki), which are too large to commit.
## Remaining (honest)
Of the 11 ELP backfill targets, 4 are done (es fr de la). The other 7 have **no
full-lexicon engine** yet — cannot be generated honestly without engine work:
- **ru**: only a 110-entry curated Slavic subset exists; full `rus.unimorph`
present but no `morphology_ru_full` productive loader. Needs a full Russian
morphology module (like the Romance ones) before vocab generation.
- **ja / ko / zh**: validated demo engines (~66-104 hardcoded words) in
`lang-poetry-experiment`, Python only. Agglutinative (ja/ko) + isolating (zh)
need `.el` engine ports + full-lexicon wiring (ja: jpn_unimorph; zh: CC-CEDICT).
- **ar / he (Semitic)**: template engines (16 AR / 8 HE patterns, ~6 roots) in
`~/semitic_engine`, Python only. Root-and-pattern; full UniMorph ara/heb
present but used only for validation. Needs productive root lexicon + `.el` port.
- **hi (Hindi), fi (Finnish), sw (Swahili)**: `morphology-<code>.el` exists in
ELP but there is NO scattered prototype and NO downloaded data for these —
full-lexicon collection (UniMorph/kaikki) + generator still to do.
De/nl/sv Germanic and it/ro/ca/pt Romance verb coverage note: German verbs here
are the ~6.6k caches carry; the it/ro/ca/pt bonus languages have full vocab but
**no `morphology-<code>.el` in ELP yet** (Python realizer exists; `.el` port is
the remaining engine work).
Construction coverage (separate from lexicon): French realizer was ~55%,
Semitic ~3% in the prototypes — full construction coverage remains its own task.
-5
View File
@@ -80,11 +80,6 @@ build {
"src/grammar.el",
"src/realizer.el",
"src/semantics.el",
"src/comprehend.el",
"src/propositions.el",
"src/multilingual.el",
"src/self_region.el",
"src/dialogue.el",
"src/elp.el",
]
}
File diff suppressed because it is too large Load Diff
-16
View File
@@ -1,16 +0,0 @@
// comprehend.elh — public surface of the ELP comprehension front-end.
// text → meaning-spec (the input half of the ELP; inverse of the realizer).
extern fn parse_spec(text: String) -> [String]
extern fn parse_spec_lang(text: String, lang: String) -> [String]
extern fn parse_json(text: String) -> String
extern fn parse_json_lang(text: String, lang: String) -> String
// Analysis primitives (invertible morphology + deterministic grammar helpers):
extern fn cp_tokenize(text: String) -> [String]
extern fn cp_pron_concept(w: String) -> String
extern fn cp_is_negation(w: String) -> Bool
extern fn cp_is_neg_adverb(w: String) -> Bool
extern fn cp_irr2(surface: String) -> [String]
extern fn cp_reg_verb(w: String) -> [String]
extern fn cp_analyze_verb(surface: String) -> [String]
extern fn cp_verb_start(toks: [String], end: Int) -> Int
extern fn cp_subord_start(toks: [String], n: Int) -> Int
-287
View File
@@ -1,287 +0,0 @@
// dialogue.el SUMMON-THROUGH-SELF, native el. Port of dialogue.py's core.
//
// THE WHOLE DIALOGUE IS ONE OPERATION. A fact is never merely *fetched*: the
// query is PROJECTED into the engram's self + memory geometry, LANDS in a region,
// and the reply is READ OUT / the region MATERIALIZED from wherever it landed.
//
// project(query) -> land on a region -> read out from that region
//
// lands in the SELF region -> grounded identity/presence, read out of
// the real self nodes (self_region.el)
// lands on a memory NEIGHBORHOOD -> MATERIALIZE it: walk the neighborhood
// (engram_neighbors_json) and read out the
// region's connected members
// lands nowhere close -> HONEST ABSENCE (an empty region, not a
// fabricated answer, not an error)
//
// CRITICAL INVARIANTS (enforced structurally, not by convention):
// * ONE operation there is NO intent classifier and NO separate
// fact-retrieval branch. Identity is nearest-region proximity, not a switch.
// * MATERIALIZE by walking the neighborhood, never by fetching top-props.
// * HONEST ABSENCE when the region is thin.
// * NEGATION is SACRED: the readout is the stored prose VERBATIM, so a negated
// memory stays negated we never paraphrase a polarity away.
// * NO ECHO: the old "I noted that X. That relates to Y." template is gone.
// The summon path materializes or honestly declines it never echoes.
// * DIRECTIVE OVERRIDE: a meta-directive ("answer in English") overrides the
// reply language while the content language is still auto-detected.
//
// Depends on: comprehend (parse_spec_lang, cp_tokenize), multilingual (ml_detect,
// ml_tr, ml_term), propositions (prop_split_sentences), self_region
// (sr_available, sr_readout), the engram + json runtime builtins.
// directive override
// Return [target_lang, content]. target_lang is "" when no directive is present.
// A directive names an output language; we strip it and keep the remaining text
// as the content (whose OWN language is still auto-detected downstream).
fn dlg_dir_hit(low: String, phrase: String) -> Bool {
return str_contains(low, phrase)
}
fn dlg_parse_directive(text: String) -> [String] {
let low: String = str_to_lower(text)
let lang: String = ""
let phrase: String = ""
// English target
if dlg_dir_hit(low, "in english") { let lang = "en"; let phrase = "in english" }
if dlg_dir_hit(low, "em inglês") { let lang = "en"; let phrase = "em inglês" }
if dlg_dir_hit(low, "em ingles") { let lang = "en"; let phrase = "em ingles" }
if dlg_dir_hit(low, "en inglés") { let lang = "en"; let phrase = "en inglés" }
// Portuguese target
if dlg_dir_hit(low, "in portuguese") { let lang = "pt"; let phrase = "in portuguese" }
if dlg_dir_hit(low, "em português") { let lang = "pt"; let phrase = "em português" }
// Spanish target
if dlg_dir_hit(low, "in spanish") { let lang = "es"; let phrase = "in spanish" }
if dlg_dir_hit(low, "en español") { let lang = "es"; let phrase = "en español" }
// Italian target
if dlg_dir_hit(low, "in italian") { let lang = "it"; let phrase = "in italian" }
let content: String = text
if !str_eq(phrase, "") {
// strip the directive phrase (and a common "answer"/"responda" lead-in),
// leaving the real question as content.
let idx: Int = str_index_of(low, phrase)
if idx >= 0 {
let before: String = str_slice(text, 0, idx)
let after: String = str_slice(text, idx + str_len(phrase), str_len(text))
let content = str_trim(before + " " + after)
}
// trim a leading "answer"/"responda"/"reply" and stray colon/comma.
let cl: String = str_to_lower(content)
if str_starts_with(cl, "answer") { let content = str_trim(str_slice(content, 6, str_len(content))) }
if str_starts_with(cl, "responda") { let content = str_trim(str_slice(content, 8, str_len(content))) }
if str_starts_with(cl, "reply") { let content = str_trim(str_slice(content, 5, str_len(content))) }
if str_starts_with(content, ":") { let content = str_trim(str_slice(content, 1, str_len(content))) }
if str_starts_with(content, ",") { let content = str_trim(str_slice(content, 1, str_len(content))) }
}
let r: [String] = native_list_empty()
let r = native_list_append(r, lang)
let r = native_list_append(r, content)
return r
}
// identity landing (a region proximity, not a classifier switch)
// The query lands in the SELF region when it takes an identity/presence shape.
// Cross-lingual forms are included because the engram's lexical probe is
// English-leaning. This is the SELF attractor of the single operation.
fn dlg_is_identity(content: String) -> Bool {
let low: String = str_to_lower(str_trim(content))
if str_contains(low, "who are you") { return true }
if str_contains(low, "what are you") { return true }
if str_contains(low, "who i am") { return true }
if str_contains(low, "your name") { return true }
if str_contains(low, "about yourself") { return true }
if str_contains(low, "are you conscious") { return true }
if str_contains(low, "are you there") { return true }
// cross-lingual identity question-forms
if str_contains(low, "quem é você") { return true }
if str_contains(low, "quem es voce") { return true }
if str_contains(low, "quién eres") { return true }
if str_contains(low, "quien eres") { return true }
if str_contains(low, "chi sei") { return true }
if str_contains(low, "qui es-tu") { return true }
if str_contains(low, "wer bist du") { return true }
return false
}
// readout helpers
fn dlg_first_sentence(content: String) -> String {
let sents: [String] = prop_split_sentences(content)
let n: Int = native_list_len(sents)
let i: Int = 0
while i < n {
let s: String = str_trim(native_list_get(sents, i))
// drop a leading markdown heading marker for a clean read-out line
if str_starts_with(s, "# ") { let s = str_trim(str_slice(s, 2, str_len(s))) }
if str_len(s) > 0 { return s }
let i = i + 1
}
return str_trim(content)
}
// strip trailing/leading punctuation from a token.
fn dlg_clean_tok(w: String) -> String {
let s: String = str_trim(w)
let s = str_strip_suffix(s, ".")
let s = str_strip_suffix(s, ",")
let s = str_strip_suffix(s, "?")
let s = str_strip_suffix(s, "!")
let s = str_strip_suffix(s, ":")
let s = str_strip_suffix(s, ";")
return str_trim(s)
}
// closed-class across the supported languages (union) a word we must NOT treat
// as a retrieval topic. Also drops the meta verbs of a request ("tell", "prove",
// "show") so the TOPIC, not the speech act, is what projects into memory.
fn dlg_is_stop(w: String) -> Bool {
if ml_stop_en(w) { return true }
if ml_stop_es(w) { return true }
if ml_stop_pt(w) { return true }
if ml_stop_it(w) { return true }
if str_eq(w, "tell") { return true }
if str_eq(w, "show") { return true }
if str_eq(w, "about") { return true }
if str_eq(w, "sobre") { return true }
if str_eq(w, "acerca") { return true }
return false
}
// The CONTENT TERMS the query projects into memory: content words only, cleaned,
// cross-lingually mapped to the engram's English vocabulary, 3 chars. This is
// the geometry probe the speech-act verbs and function words are stripped so a
// PP topic ("tell me ABOUT Lisbon") projects on "lisbon", not "tell"/"me".
fn dlg_content_terms(content: String, lang: String) -> [String] {
let toks: [String] = cp_tokenize(content)
let n: Int = native_list_len(toks)
let out: [String] = native_list_empty()
let i: Int = 0
while i < n {
let w: String = str_to_lower(dlg_clean_tok(native_list_get(toks, i)))
if str_len(w) >= 3 {
if !dlg_is_stop(w) {
let out = native_list_append(out, ml_term(w, lang))
}
}
let i = i + 1
}
return out
}
// Does this landed node lexically overlap the query's content terms? This is the
// RELEVANCE FLOOR: activation always returns the store's most salient nodes, so
// without this a query about nothing would "land" on the self/top node. A node
// that shares no content term with the query is "nowhere close" -> honest absence.
fn dlg_node_matches(node: String, terms: [String]) -> Bool {
let hay: String = str_to_lower(json_get_string(node, "content") + " " + json_get_string(node, "label"))
let n: Int = native_list_len(terms)
let i: Int = 0
while i < n {
let t: String = native_list_get(terms, i)
if str_len(t) >= 3 {
if str_contains(hay, t) { return true }
}
let i = i + 1
}
return false
}
// MATERIALIZE the landed region: read out the landed fact, then WALK the
// neighborhood and read out its connected members (real edges, not top-props).
fn dlg_materialize(top_node: String, reply_lang: String) -> String {
let id: String = json_get_string(top_node, "id")
let content: String = json_get_string(top_node, "content")
let lead: String = dlg_first_sentence(content)
let nb: String = engram_neighbors_json(id, 2, "both")
let m: Int = json_array_len(nb)
let parts: [String] = native_list_empty()
let parts = native_list_append(parts, lead)
let added: Int = 0
let i: Int = 0
while i < m {
if added < 3 {
let rec: String = json_array_get(nb, i)
let node: String = json_get_raw(rec, "node")
let nc: String = json_get_string(node, "content")
if !str_eq(nc, "") {
let sent: String = dlg_first_sentence(nc)
if !str_eq(sent, "") {
let parts = native_list_append(parts, sent)
let added = added + 1
}
}
}
let i = i + 1
}
// The readout is the region's OWN prose, verbatim negation SACRED, no echo.
return str_join(parts, " ")
}
// THE single operation
fn dlg_respond(text: String) -> String {
// directive override: reply language may differ from content language.
let dir: [String] = dlg_parse_directive(text)
let target_lang: String = native_list_get(dir, 0)
let content: String = native_list_get(dir, 1)
let content_lang: String = ml_detect(content)
let reply_lang: String = content_lang
if !str_eq(target_lang, "") { let reply_lang = target_lang }
// comprehend the content (SACRED polarity carried in the spec).
let spec: [String] = parse_spec_lang(content, content_lang)
// PROJECT + LAND: SELF region
// Identity/presence shape lands in the self region; read out the REAL self
// nodes (self_region.el), never a template. Same single operation this is
// just the self attractor winning the landing.
if dlg_is_identity(content) {
if sr_available() {
// read out the REAL self nodes when replying in their own language
// (the soul's prose is English); for another reply language we cannot
// translate real content without an LLM, so we answer with the
// localized SACRED identity anchor honest, in-language, no fabrication.
if str_eq(reply_lang, "en") { return sr_readout("en") }
return ml_tr("identity", reply_lang)
}
// self region thin honest localized identity (logged fallback shape).
return ml_tr("identity", reply_lang)
}
// PROJECT into MEMORY geometry
let terms: [String] = dlg_content_terms(content, content_lang)
let qterm: String = str_join(terms, " ")
let act: String = engram_activate_json(qterm, 12)
let n: Int = json_array_len(act)
// LAND: the highest-activation node that ACTUALLY overlaps the query's
// content terms (the relevance floor). Activation always returns the most
// salient nodes, so we walk the ranked list and take the first that is
// genuinely "close"; if none is, the query landed nowhere. ───────────────
let landing: String = ""
let i: Int = 0
while i < n {
if str_eq(landing, "") {
let rec: String = json_array_get(act, i)
let node: String = json_get_raw(rec, "node")
if dlg_node_matches(node, terms) {
let landing = node
}
}
let i = i + 1
}
// HONEST ABSENCE: nothing close an empty region, not a fabricated answer,
// not an "I noted that" echo.
if str_eq(landing, "") {
return ml_tr("no_memory", reply_lang)
}
// MATERIALIZE the landing by WALKING its neighborhood.
return dlg_materialize(landing, reply_lang)
}
-13
View File
@@ -63,9 +63,6 @@ import "morphology-cop.el"
import "grammar.el"
import "realizer.el"
import "semantics.el"
// Comprehension front-end (input half: text meaning-spec)
import "comprehend.el"
//
// Entry points:
//
@@ -120,9 +117,6 @@ fn build_form_from_json(semantic_form_json: String, lang_code: String) -> [Strin
let location: String = sem_get(semantic_form_json, "location")
let tense: String = sem_get(semantic_form_json, "tense")
let aspect: String = sem_get(semantic_form_json, "aspect")
let polarity: String = sem_get(semantic_form_json, "polarity")
let neg_word: String = sem_get(semantic_form_json, "neg_word")
let iobj: String = sem_get(semantic_form_json, "iobj")
let form: [String] = native_list_empty()
let form = native_list_append(form, "intent")
@@ -133,19 +127,12 @@ fn build_form_from_json(semantic_form_json: String, lang_code: String) -> [Strin
let form = native_list_append(form, predicate)
let form = native_list_append(form, "patient")
let form = native_list_append(form, patient)
let form = native_list_append(form, "iobj")
let form = native_list_append(form, iobj)
let form = native_list_append(form, "location")
let form = native_list_append(form, location)
let form = native_list_append(form, "tense")
let form = native_list_append(form, tense)
let form = native_list_append(form, "aspect")
let form = native_list_append(form, aspect)
// SACRED: polarity crosses the JSON boundary and is never inferred away.
let form = native_list_append(form, "polarity")
let form = native_list_append(form, polarity)
let form = native_list_append(form, "neg_word")
let form = native_list_append(form, neg_word)
let form = native_list_append(form, "lang")
let form = native_list_append(form, lang_code)
-72
View File
@@ -1,72 +0,0 @@
;;; lang_profile_ca.el — Catalan language profile for ELP.
;;; Mirrors lang_profile_it / _es / _pt; keys the realizer's construction switches.
;;; Catalan is the CLOSEST Romance sibling to the shared engine (~85% conceptual
;;; reuse). The deltas: PRONOMS FEBLES with four position allomorphs, l'-elision,
;;; del/al/pel contractions, the periphrastic preterite (vaig+INF), and NO
;;; essere/avere split (perfect aux is always HAVER; ser/estar is only the copula).
(lang_profile_ca
(language "Catalan")
(iso639 "ca")
(family "Romance")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop yes) ; null subjects default; overt pronoun = emphatic
(obligatory-subject no)
(grammatical-gender yes) ; m/f; full NP agreement (art + adj + participle)
(do-support no)
(subject-aux-inversion no) ; yes/no Q = declarative order + '?'; no inversion
(article-selection "el/la/l'/els/les ; un/una/uns/unes") ; l'-ELISION:
; el/la -> l' before vowel or (silent) h, glued to
; the next word (l'home, l'illa); de -> d' before vowel
(article-drives-contraction yes) ; article choice feeds prep+article contraction
(adjective-position "postnominal-default + small prenominal class") ; bo/bon,
; mal, gran, nou, vell, primer, molt... prenominal
(question-punct plain) ; ? and ! only (no inverted ¿ ¡)
;; ── MANDATORY prep+article contractions ────────────────────────────────
(contractions ((de el del) (de els dels)
(a el al) (a els als)
(per el pel) (per els pels)))
(contraction-mandatory yes) ; *de el -> del obligatory
(contraction-blocked-before-elision yes) ; de l'home / a l'home (NO *del home)
;; ── clitic system: PRONOMS FEBLES (the headline delta) ──────────────────
(clitics yes)
(clitic-allomorphy four-position) ; per pronoun, form varies by position+onset:
; reinforced (em, et, el) proclitic before a consonant
; elided (m', t', l', n') proclitic before a vowel/h
; full (-me, -lo, -li) enclitic after a consonant/-r
; reduced ('m, 't, 'l, 'ns) enclitic after a vowel
(clitic-placement ((finite proclitic) ; el veig, no m'ho dóna
(imperative-affirmative enclitic) ; dóna'm, digues-me
(imperative-negative present-subjunctive) ; no parlis (delta)
(infinitive enclitic) ; ajudar-me, veure'l
(gerund enclitic))) ; fent-ho
(clitic-combination ((me el "me'l") (te el "te'l") (se el "se'l")
(me la "me la") (me en "me'n")
(li el "l'hi") (li en "n'hi"))) ; dative+accusative clusters
(clitic-particles (hi en ho)) ; locative hi, partitive/genitive en, neuter ho
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "person+number (6-way)")
(tenses (present imperfet preterit-simple perifrastic-preterit futur
condicional subjuntiu-present subjuntiu-imperfet imperatiu))
(periphrastic-preterite "vaig/vas/va/vam/vau/van + INFINITIVE") ; << hallmark CA
; (vaig cantar = 'I sang'); coexists w/ synthetic pret.
(compound-past "pretèrit perfet = haver(present) + participle")
(perfect-aux "HAVER only") ; << NO essere/avere split (simpler than IT)
(participle-agreement ((haver preceding-acc-clitic))) ; les he vistes; else invariable
(progressive-aux "estar + gerundi")
(copula "ser / estar") ; ser: identity/essential/origin; estar:
; location + transient state (estic cansat, és a casa)
(passive-aux "ser (+ per-agent)")
(future inflectional) ; cantaré, serà
(comparative "més/menys ADJ que")
;; ── SACRED safety bar (shared with es/pt/it/en) ────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
(negation "no (preverbal) + optional 'pas' + concord") ; no...res/
; ningú/mai/cap/gens/enlloc
(negative-concord yes) ; preverbal negative subject (ningú) keeps 'no'
(neg-reinforcer pas)) ; optional (no ho faré pas)
-41
View File
@@ -1,41 +0,0 @@
;;; lang_profile_de.el — German language profile for ELP.
;;; Mirrors lang_profile_en / lang_profile_es. Keys the realizer's construction
;;; switches. German is the largest Germanic delta from the EN engine: V2 word
;;; order, four morphological cases, and separable-prefix verbs.
(lang_profile_de
(language "German")
(iso639 "de")
(family "Germanic")
(neighbor-base "en") ; realized by extending the English (Germanic) engine
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop no) ; obligatory subject in finite clauses
(obligatory-subject yes)
(grammatical-gender (m f n)) ; three genders; drives article + adj declension
(case-system (nom acc dat gen)) ; four cases on articles/adjs/nouns
(word-order V2) ; finite verb 2nd in main clause
(subordinate-order verb-final) ; "..., dass er den Hund SIEHT."
(separable-verbs yes) ; aufstehen -> "steht ... auf"; ppart "aufgestanden"
(do-support no) ; German negates/questions the finite verb directly
(subject-verb-inversion yes) ; yes/no Q fronts finite verb; wh-Q fills Vorfeld
(article-selection "der/die/das + ein/kein") ; declined by case x gender x number
(adjective-position prenominal)
(adjective-declension (strong weak mixed)) ; chosen by the determiner type
(noun-capitalization yes)
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "person-and-number") ; full present/past paradigm
(auxiliary-order (modal tense-aux perfect passive main))
(perfect-aux (haben sein)) ; sein for intransitive motion/change verbs
(passive-aux "werden")
(future "werden + infinitive")
(comparative "synthetic (-er / -st, with umlaut)")
;; ── negation ───────────────────────────────────────────────────────────
(negation-markers (nicht kein)) ; kein- negates an indefinite NP; nicht else
(negation-faithful yes) ; SACRED: polarity never dropped/inverted -> FLAG
;; ── lexicon provenance ─────────────────────────────────────────────────
(lexicon-source "UniMorph deu (primary) + kaikki.org German (gender override)")
(lexicon-license "CC-BY-SA 3.0 / GFDL"))
-41
View File
@@ -1,41 +0,0 @@
;;; lang_profile_en.el — English language profile for ELP.
;;; Mirrors lang_profile_es / lang_profile_pt; keys the realizer's construction
;;; switches. English is typologically distinct from the Romance builds, so the
;;; flags differ where the grammar differs.
(lang_profile_en
(language "English")
(iso639 "en")
(family "Germanic")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop no) ; OBLIGATORY subjects — missing subject is FLAGGED
(obligatory-subject yes)
(grammatical-gender no) ; natural gender only (he/she/it), no NP agreement
(do-support yes) ; negation & questions of lexical verbs insert do/does/did
(subject-aux-inversion yes) ; yes/no + non-subject wh questions invert the operator
(article-selection "a/an/the") ; a/an resolved PHONOLOGICALLY (an hour, a university)
(adjective-position prenominal) ; attributive adjectives precede the noun; invariant
(has-tag-questions yes) ; "...doesn't he?" — operator + reversed polarity
(has-there-existential yes) ; "there is/are/have been ..."
(possessive-clitic "'s") ; saxon genitive; plural in -s -> bare apostrophe
(question-punct plain) ; ? and ! only (no inverted marks)
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "3sg-present-only") ; only 3sg present -s (+ suppletive be)
(auxiliary-order (modal perfect progressive passive main))
(perfect-aux "have") ; have + past participle
(progressive-aux "be") ; be + present participle
(passive-aux "be") ; be + past participle (+ by-agent)
(future "will + base") ; no inflectional future
(comparative "synthetic-or-periphrastic") ; -er/-est vs more/most by syllables
;; ── SACRED safety bar (shared with es/pt) ──────────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
;; ── DIALECT overlay (post-realization, one core -> US/UK/AU) ────────────
(dialect US) ; default; profile field switches the overlay
(dialects (US UK AU))
(dialect-canonical US) ; core is authored in US orthography
(dialect-overlay "dialect_en.to_dialect") ; orthography + lexis + grammar prefs
(dialect-covers (spelling lexis collective-agreement gotten/got)))
-45
View File
@@ -1,45 +0,0 @@
;;; lang_profile_es.el — Spanish language profile for ELP.
;;; Keys the realizer's construction switches. Mirrors lang_profile_en / _pt.
(lang_profile_es
(language "Spanish")
(iso639 "es")
(family "Romance")
;; -- core typology flags -------------------------------------------------
(pro-drop yes) ; subjects routinely dropped; agreement carries person
(obligatory-subject no)
(grammatical-gender yes) ; m/f on every noun; article+adjective AGREE
(gender-source lexicon); REAL per-noun gender from UniMorph — NOT a heuristic
(do-support no)
(subject-aux-inversion no) ; questions by intonation/punctuation, not inversion
(question-strategy intonation)
(article-selection "el/la/los/las un/una/unos/unas")
(stressed-a-rule yes) ; fem sg noun in stressed a-/ha- takes el/un (el agua)
(adjective-position postnominal) ; default post; a few prenominal + apocope
(adjective-agreement "gender+number")
(question-punct inverted) ; opening ¿ ¡ required
;; -- MANDATORY CONTRACTIONS (coordinator quality bar) --------------------
(contractions ((de el "del") (a el "al")))
(contraction-mandatory yes) ; 'de el'/'a el' MUST surface as del/al
;; -- verb / aspect system ------------------------------------------------
(verb-classes (ar er ir))
(tenses (present preterite imperfect future conditional))
(moods (ind sbjv imp))
(finite-agreement "person+number (6 slots)")
(perfect-aux "haber") ; haber + past participle (invariant -o)
(progressive-aux "estar") ; estar + gerund
(passive-aux "ser") ; ser + participle (agrees) + por-agent
(copula-split "ser/estar") ; permanent vs stage-level
(future "infinitive + é/ás/á/emos/éis/án")
;; -- clitics / government ------------------------------------------------
(object-clitics yes) ; me te lo la le nos os los las; proclisis/enclisis
(clitic-order "se II I III (le+lo -> se lo)")
(enclisis "imperative/infinitive/gerund + accent repair (dá+me+lo->dámelo)")
(verb-prep-government yes) ; verbs select prep (protestar+contra, escapar+de)
;; -- SACRED safety bar (shared with en/pt) -------------------------------
(negation-faithful yes)) ; polarity never dropped/inverted; unplaceable -> FLAG
-74
View File
@@ -1,74 +0,0 @@
;;; lang_profile_fr.el — French language profile for ELP.
;;; Mirrors lang_profile_it / lang_profile_es; keys the realizer's construction
;;; switches. French is a Romance sibling (~54% of the realizer code and the whole
;;; clause-engine architecture reused), but carries the family's biggest surface
;;; deltas: NOT pro-drop, DISCONTINUOUS negation, and an orthography/phonology
;;; mismatch (elision, liaison) that makes exact-match genuinely hard.
(lang_profile_fr
(language "French")
(iso639 "fr")
(family "Romance")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop no) ; << French-specific: subject clitic OBLIGATORY
(obligatory-subject yes) ; je/tu/il/elle/nous/vous/ils/elles always overt
(grammatical-gender yes) ; m/f; full NP agreement (art + adj + participle)
(do-support no)
(subject-aux-inversion optional) ; est-ce que (default) OR clitic inversion (vas-tu)
(article-selection "le/la/l'/les ; un/une/des ; PARTITIVE du/de la/de l'/des")
(article-drives-contraction yes) ; à+le=au, de+le=du feed off article choice
(adjective-position "postnominal-default + prenominal-BAGS") ; beau/bon/grand/
; petit/jeune/vieux/nouveau + ordinals prenominal
; (beau->bel, nouveau->nouvel, vieux->vieil / vowel)
(question-punct "space-before") ; French typography: ' ?' ' !' (no ¿¡)
;; ── elision (orthography/phonology mismatch — French-specific) ──────────
(elision ((le l') (la l') (je j') (ne n') (de d') (que qu')
(me m') (te t') (se s') (ce c'))) ; before vowel / h-muet
(elision-h-muet yes) ; l'homme, l'hôpital (h-aspiré exception list kept)
(liaison noted-not-modeled) ; phonological, not written in surface
;; ── MANDATORY prep+article contractions ────────────────────────────────
(contractions ((à le au) (à les aux) (de le du) (de les des)))
(contraction-mandatory yes) ; *à le -> au obligatory; à la / à l' uncontracted
(partitive ((m-sg du) (f-sg "de la") (vowel "de l'") (pl des)))
(partitive-under-neg "de") ; << gap in current build: 'ne … pas de pain'
;; ── clitic system ──────────────────────────────────────────────────────
(clitics yes)
(clitic-order (me te se nous vous | le la les | lui leur | y | en))
(clitic-placement ((finite proclitic) ; je le lui donne
(imperative-affirmative enclitic-hyphen) ; donne-le-moi
(imperative-negative "ne+proclitic+verb+pas") ; ne le donne pas
(infinitive enclitic))) ; PARTIAL: clitic-climbing
; onto infinitive under modal
(clitic-imperative-shift ((me moi) (te toi))) ; final me/te -> moi/toi (donne-moi)
(clitic-particles (y en)) ; locative y, partitive/genitive en
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "person+number (written; many homophones)")
(tenses (présent imparfait passé-simple futur conditionnel
subjonctif-présent subjonctif-imparfait impératif))
(compound-past "passé-composé = aux(present) + participe passé")
(perfect-aux "être/avoir (LEXICAL selection)") ; << French-specific
(etre-aux-class "intransitive motion/change (aller venir arriver partir
entrer sortir monter descendre naître mourir rester
tomber retourner passer devenir revenir rentrer) + ALL
pronominal verbs")
(participle-agreement ((être subject) ; elle est allée / elles venues
(avoir preceding-direct-object))) ; je les ai vus
(progressive "être en train de + infinitif") ; no dedicated aux
(copula "être (single; no ser/estar, no essere/stare)")
(passive-aux "être (+ par-agent)")
(future inflectional) ; parlera, sera
(comparative "plus/moins ADJ que")
(superlative "le/la plus ADJ (de …)") ; PARTIAL word-order in build
;; ── SACRED safety bar (shared with es/pt/it/en) ────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
(negation "DISCONTINUOUS: ne (preverbal) … pas/jamais/rien/personne/
plus/guère/que (postverbal)") ; << biggest structural delta
(negation-ne-elides yes) ; ne -> n' before vowel (n'ai pas vu)
(negation-passe-composé "ne + aux + pas + participe") ; n'ai pas vu
(negative-concord partial)) ; personne/rien as arguments post-participle
-70
View File
@@ -1,70 +0,0 @@
;;; lang_profile_it.el — Italian language profile for ELP.
;;; Mirrors lang_profile_es / lang_profile_pt; keys the realizer's construction
;;; switches. Italian is a Romance sibling, so ~85% of the flags match ES/PT; the
;;; essere/avere auxiliary split and phonological article selection are the deltas.
(lang_profile_it
(language "Italian")
(iso639 "it")
(family "Romance")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop yes) ; null subjects default; overt pronoun = emphatic
(obligatory-subject no)
(grammatical-gender yes) ; m/f; full NP agreement (art + adj + participle)
(do-support no)
(subject-aux-inversion no) ; yes/no Q = declarative order + '?'; no inversion
(article-selection "il/lo/l'/i/gli + la/l'/le ; un/uno/un'/una") ; PHONOLOGICAL:
; lo/gli/uno before s+cons, z, gn, ps, pn, x, y, i+V;
; l'/un' before a vowel (elision, glued to next word)
(article-drives-contraction yes) ; article choice feeds the prep+art contraction
(adjective-position "postnominal-default + prenominal-class") ; bello/buono/grande
; /nuovo/vecchio/primo... prenominal (with apocope)
(question-punct plain) ; ? and ! only (no inverted ¿ ¡)
;; ── MANDATORY prep+article contractions ────────────────────────────────
(contractions ((di il del) (di lo dello) (di la della) (di i dei)
(di gli degli) (di le delle) (di l' dell')
(a il al) (a lo allo) (a la alla) (a i ai) (a gli agli)
(a le alle) (a l' all')
(da il dal) (da la dalla) (da gli dagli) (da l' dall')
(in il nel) (in la nella) (in gli negli) (in l' nell')
(su il sul) (su la sulla) (su gli sugli) (su l' sull')))
(contraction-mandatory yes) ; *di il -> del is obligatory, never uncontracted
(prep-no-contract (per tra fra)) ; per la strada (NOT *perla)
;; ── clitic system ──────────────────────────────────────────────────────
(clitics yes)
(clitic-placement ((finite proclitic) ; lo vedo, non me lo dà
(imperative-affirmative enclitic) ; dammelo, guardalo
(imperative-negative-tu non+infinitive) ; non parlare / non lo fare
(infinitive enclitic) ; vederlo, aiutarmi (drop -e)
(gerund enclitic))) ; dandolo
(clitic-combination ((mi lo "me lo") (ti lo "te lo") (ci lo "ce lo")
(vi lo "ve lo") (si lo "se lo")
(gli lo "glielo") (le lo "glielo"))) ; glielo = ONE word
(clitic-particles (ci ne)) ; locative ci, partitive ne
(raddoppiamento (da fa di va sta)) ; monosyllabic imper double clitic: dammelo
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "person+number (6-way)")
(tenses (presente imperfetto passato-remoto futuro condizionale
congiuntivo-presente congiuntivo-imperfetto imperativo))
(compound-past "passato-prossimo = aux(present) + participle")
(perfect-aux "essere/avere (LEXICAL selection)") ; << Italian-specific
(essere-aux-class unaccusative) ; motion/change-of-state/copular/pronominal
; (andare venire nascere morire diventare piacere
; + ALL reflexives) -> essere
(participle-agreement ((essere subject) ; è andata / sono arrivati
(avere preceding-acc-clitic))) ; li ho visti
(progressive-aux "stare + gerundio") ; sto parlando
(copula "essere (default) / stare (state: sto bene)")
(passive-aux "essere / venire (+ da-agent)")
(future inflectional) ; parlerò, sarà
(comparative "più/meno ADJ di")
;; ── SACRED safety bar (shared with es/pt/en) ───────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
(negation "non (preverbal) + concord") ; non...niente/nessuno/mai/più
(negative-concord yes) ; preverbal negative word (nessuno/niente) suppresses non
(neg-adverb-position between-aux-and-participle)) ; non ho MAI visto
-30
View File
@@ -1,30 +0,0 @@
;;; lang_profile_la.el — Latin language profile for ELP.
;;; Keys the realizer's construction switches. Companion to morphology-la.el.
(lang_profile_la
(language "Latin")
(iso639 "la")
(family "Italic")
;; -- core typology flags -------------------------------------------------
(pro-drop yes) ; person carried by verb ending; subjects dropped
(obligatory-subject no)
(grammatical-gender yes) ; m/f/n; adjective AGREES in case+gender+number
(gender-source lexicon) ; REAL per-noun gender from UniMorph lat
(articles none) ; Latin has no articles
(case-system yes) ; NOM GEN DAT ACC ABL VOC (+ rare LOC)
(cases (nom gen dat acc abl voc))
(word-order "SOV (default; free order, case-marked)")
(adjective-position "either (case agreement carries the link)")
(adjective-agreement "case+gender+number")
;; -- verb / aspect system ------------------------------------------------
(verb-classes (1 2 3 3io 4)) ; four conjugations + i-stem 3rd
(tenses (present imperfect future perfect pluperfect futureperfect))
(moods (indicative subjunctive imperative infinitive))
(voices (active passive))
(finite-agreement "person+number (6 slots)")
(citation "principal parts: pres-1sg / pres-inf / perf-participle")
;; -- SACRED safety bar ---------------------------------------------------
(negation-faithful yes)) ; polarity never dropped/inverted
-40
View File
@@ -1,40 +0,0 @@
;;; lang_profile_pt.el — Portuguese language profile for ELP.
;;; Keys the realizer's construction switches. Mirrors lang_profile_es.
(lang_profile_pt
(language "Portuguese")
(iso639 "pt")
(family "Romance")
;; -- core typology flags -------------------------------------------------
(pro-drop yes) ; subjects routinely dropped; agreement carries person
(obligatory-subject no)
(grammatical-gender yes) ; m/f on every noun; article+adjective AGREE
(gender-source lexicon) ; REAL per-noun gender from UniMorph por / kaikki
(do-support no)
(subject-aux-inversion no)
(question-strategy intonation)
(article-selection "o/a/os/as um/uma/uns/umas")
(adjective-position postnominal)
(adjective-agreement "gender+number")
;; -- MANDATORY CONTRACTIONS (prep + article) -----------------------------
(contractions ((de o "do") (de a "da") (em o "no") (em a "na")
(a o "ao") (a a "à") (por o "pelo") (por a "pela")))
(contraction-mandatory yes)
;; -- verb / aspect system ------------------------------------------------
(verb-classes (ar er ir))
(tenses (present preterite imperfect future conditional))
(moods (ind sbjv imp))
(finite-agreement "person+number (6 slots)")
(perfect-aux "ter") ; ter + past participle
(copula-split "ser/estar")
(personal-infinitive yes) ; distinctive PT inflected infinitive
;; -- clitics / government ------------------------------------------------
(object-clitics yes) ; mesoclisis/enclisis/proclisis by context
(verb-prep-government yes)
;; -- SACRED safety bar ---------------------------------------------------
(negation-faithful yes))
-71
View File
@@ -1,71 +0,0 @@
;;; lang_profile_ro.el — Romanian language profile for ELP.
;;; Romanian is the BIG typological delta of the Romance family. The verb/clause
;;; engine and the SACRED negation contract mirror the ES/PT/IT core, but the
;;; NOMINAL system is genuinely new: a SUFFIXED definite article, preserved CASE,
;;; a NEUTER gender, and a VOCATIVE. Those flags mark where the shared engine was
;;; extended rather than reused.
(lang_profile_ro
(language "Romanian")
(iso639 "ro")
(family "Romance (Eastern / Balkan)")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop yes) ; null subjects default; overt pronoun = emphatic
(obligatory-subject no)
(grammatical-gender yes) ; m / f / NEUTER (n)
(neuter-gender yes) ; << ROMANIAN-SPECIFIC: masc-agreeing SG, fem-agreeing PL
; (un tren nou / două trenuri noi)
(do-support no)
(subject-aux-inversion no) ; yes/no Q = declarative order + '?'
(question-punct plain) ; ? and ! only
;; ── SUFFIXED DEFINITE ARTICLE (the headline engine extension) ───────────
(definite-article suffixed) ; << UNIQUE IN ROMANCE: enclitic on the noun
(definite-forms ((m/n sg "-ul / -le / -l : om->omul, câine->câinele, codru->codrul")
(f sg "-a / -ea / -ua : casă->casa, carte->cartea, stea->steaua")
(m pl "-i : oameni->oamenii")
(f/n pl "-le : case->casele, trenuri->trenurile")))
(article-host ((no-prenom-adj noun) ; omul bun
(prenom-adj adjective))) ; bunul om (adj carries the article)
(indefinite-article ((m/n "un") (f "o") (pl "niște") (gen/dat-pl "unor")))
;; ── CASE (preserved; NOM/ACC vs GEN/DAT) ────────────────────────────────
(case (nom/acc gen/dat vocative)) ; << ROMANIAN-SPECIFIC
(case-syncretism "nom=acc ; gen=dat")
(genitive-marking "gen/dat definite: -lui (m/n), -ei/-i (f), -lor (pl)")
(genitival-article ((m sg "al") (f sg "a") (m pl "ai") (f/n pl "ale"))) ; o carte a lui
(possession "definite-head + gen/dat possessor: casa băiatului")
(vocative ((m sg "-ule/-e : omule, băiete") (f sg "-o : Mario, fato")
(pl "-lor")))
;; ── verb / aspect system ────────────────────────────────────────────────
(finite-agreement "person+number (6-way)")
(tenses (prezent imperfect perfect-simplu conjunctiv-prezent
imperativ (periphrastic: perfect-compus viitor conditional)))
(compound-past "perfectul compus = a-avea-clitic + INVARIABLE participle")
(perfect-aux "a avea (am/ai/a/am/ați/au) — ONE auxiliary for ALL verbs")
(perfect-aux-split no) ; << SIMPLER than Italian: no essere/avere selection
(participle-agreement none) ; invariable in the perfect compus (agrees only as
; an adjective / in the passive)
(future "voi/vei/va/vom/veți/vor + infinitive (viitor literar)")
(conditional "aș/ai/ar/am/ați/ar + infinitive")
(subjunctive "conjunctiv: particle 'să' + subjunctive present")
(modal-complement "modal + să + subjunctive (vreau să merg, poți să ajuți)")
(copula "a fi")
(passive "a fi + participle (participle AGREES like an adjective)")
(comparative "mai / mai puțin ADJ decât")
;; ── clitic system (partial — see honest gaps) ───────────────────────────
(clitics yes)
(clitic-set ((acc te îl o ne îi le) (dat îmi îți îi ne le)
(refl te se ne se)))
(clitic-placement ((finite proclitic) ; îmi place, o văd
(perfect-compus elision) ; << m-am, l-am, i-am (PARTIAL)
(imperative-affirmative enclitic))) ; dă-mi (PARTIAL)
;; ── SACRED safety bar (shared with es/pt/it/en) ─────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
(negation "nu (single preverbal marker) + concord")
(negative-concord yes) ; nu … nimic / nimeni / niciodată / niciun
(negative-imperative "nu + INFINITIVE : nu pleca! (KNOWN GAP: uses imperative stem)"))
-1
View File
@@ -250,7 +250,6 @@ fn en_irregular_verb(base: String) -> [String] {
if str_eq(base, "cut") { let r: [String] = ["cut", "cuts", "cut", "cut", "cutting"]; return r }
if str_eq(base, "set") { let r: [String] = ["set", "sets", "set", "set", "setting"]; return r }
if str_eq(base, "hit") { let r: [String] = ["hit", "hits", "hit", "hit", "hitting"]; return r }
if str_eq(base, "fight") { let r: [String] = ["fight", "fights","fought", "fought", "fighting"]; return r }
return empty
}
-280
View File
@@ -1,280 +0,0 @@
// multilingual.el - the language layer for the native-el interlocutor.
//
// Deterministic, NO generative model (ports multilingual.py):
// 1. ml_detect(text) -> ISO code (en/es/pt/it) via stopword + diacritic score
// 2. ml_tr(key, lang) -> localized fixed phrase (SACRED per-language yes/no/decline)
// 3. ml_term(w, lang) -> PT/ES content term -> EN engram equivalent
// 4. ml_translate_pred(lemma, lang) -> EN predicate lemma -> target infinitive
//
// The Python detector count-weights stopwords and diacritics; here diacritics are
// scored by PRESENCE (str_contains) rather than codepoint counting, to stay clear
// of UTF-8 index hazards in the runtime. Faithful enough to classify typical
// queries; documented simplification. Depends on: comprehend (cp_tokenize).
// 1. language detection
fn ml_stop_en(w: String) -> Bool {
if str_eq(w, "the") { return true }
if str_eq(w, "does") { return true }
if str_eq(w, "do") { return true }
if str_eq(w, "did") { return true }
if str_eq(w, "what") { return true }
if str_eq(w, "who") { return true }
if str_eq(w, "is") { return true }
if str_eq(w, "are") { return true }
if str_eq(w, "how") { return true }
if str_eq(w, "you") { return true }
if str_eq(w, "your") { return true }
if str_eq(w, "of") { return true }
if str_eq(w, "to") { return true }
if str_eq(w, "and") { return true }
if str_eq(w, "for") { return true }
if str_eq(w, "explain") { return true }
if str_eq(w, "answer") { return true }
if str_eq(w, "memory") { return true }
if str_eq(w, "with") { return true }
if str_eq(w, "not") { return true }
if str_eq(w, "store") { return true }
return false
}
fn ml_stop_es(w: String) -> Bool {
if str_eq(w, "que") { return true }
if str_eq(w, "qué") { return true }
if str_eq(w, "una") { return true }
if str_eq(w, "usted") { return true }
if str_eq(w, "su") { return true }
if str_eq(w, "cómo") { return true }
if str_eq(w, "como") { return true }
if str_eq(w, "cuál") { return true }
if str_eq(w, "quién") { return true }
if str_eq(w, "está") { return true }
if str_eq(w, "es") { return true }
if str_eq(w, "los") { return true }
if str_eq(w, "las") { return true }
if str_eq(w, "del") { return true }
if str_eq(w, "al") { return true }
if str_eq(w, "explica") { return true }
if str_eq(w, "explique") { return true }
if str_eq(w, "forma") { return true }
if str_eq(w, "con") { return true }
if str_eq(w, "memoria") { return true }
if str_eq(w, "responde") { return true }
return false
}
fn ml_stop_pt(w: String) -> Bool {
if str_eq(w, "que") { return true }
if str_eq(w, "uma") { return true }
if str_eq(w, "você") { return true }
if str_eq(w, "sua") { return true }
if str_eq(w, "seu") { return true }
if str_eq(w, "como") { return true }
if str_eq(w, "memória") { return true }
if str_eq(w, "isso") { return true }
if str_eq(w, "os") { return true }
if str_eq(w, "as") { return true }
if str_eq(w, "da") { return true }
if str_eq(w, "do") { return true }
if str_eq(w, "na") { return true }
if str_eq(w, "no") { return true }
if str_eq(w, "explica") { return true }
if str_eq(w, "forma") { return true }
if str_eq(w, "é") { return true }
if str_eq(w, "está") { return true }
if str_eq(w, "com") { return true }
if str_eq(w, "responda") { return true }
return false
}
fn ml_stop_it(w: String) -> Bool {
if str_eq(w, "che") { return true }
if str_eq(w, "una") { return true }
if str_eq(w, "come") { return true }
if str_eq(w, "della") { return true }
if str_eq(w, "gli") { return true }
if str_eq(w, "è") { return true }
if str_eq(w, "sono") { return true }
if str_eq(w, "questo") { return true }
if str_eq(w, "nel") { return true }
if str_eq(w, "di") { return true }
if str_eq(w, "il") { return true }
if str_eq(w, "cosa") { return true }
if str_eq(w, "per") { return true }
if str_eq(w, "memoria") { return true }
if str_eq(w, "spiega") { return true }
if str_eq(w, "rispondi") { return true }
return false
}
// diacritic PRESENCE score (weight 3 each; hard overrides weight 8).
fn ml_dia_score(low: String, lang: String) -> Int {
let s: Int = 0
if str_eq(lang, "pt") {
if str_contains(low, "ã") { let s = s + 3 }
if str_contains(low, "õ") { let s = s + 3 }
if str_contains(low, "ç") { let s = s + 3 }
if str_contains(low, "ê") { let s = s + 3 }
if str_contains(low, "á") { let s = s + 3 }
// hard PT markers (ã/õ almost never appear outside PT)
if str_contains(low, "ã") { let s = s + 8 }
if str_contains(low, "õ") { let s = s + 8 }
}
if str_eq(lang, "es") {
if str_contains(low, "ñ") { let s = s + 3 }
if str_contains(low, "¿") { let s = s + 3 }
if str_contains(low, "¡") { let s = s + 3 }
if str_contains(low, "á") { let s = s + 3 }
if str_contains(low, "é") { let s = s + 3 }
// hard ES markers
if str_contains(low, "ñ") { let s = s + 8 }
if str_contains(low, "¿") { let s = s + 8 }
if str_contains(low, "¡") { let s = s + 8 }
}
if str_eq(lang, "it") {
if str_contains(low, "è") { let s = s + 3 }
if str_contains(low, "ì") { let s = s + 3 }
if str_contains(low, "ò") { let s = s + 3 }
}
return s
}
fn ml_stop_score(toks: [String], lang: String) -> Int {
let n: Int = native_list_len(toks)
let s: Int = 0
let i: Int = 0
while i < n {
let w: String = native_list_get(toks, i)
if str_eq(lang, "en") { if ml_stop_en(w) { let s = s + 2 } }
if str_eq(lang, "es") { if ml_stop_es(w) { let s = s + 2 } }
if str_eq(lang, "pt") { if ml_stop_pt(w) { let s = s + 2 } }
if str_eq(lang, "it") { if ml_stop_it(w) { let s = s + 2 } }
let i = i + 1
}
return s
}
fn ml_detect(text: String) -> String {
if str_eq(text, "") { return "en" }
let low: String = str_to_lower(text)
let toks: [String] = cp_tokenize(text)
// NOTE: el's overloaded `+` mis-compiles two chained function-call Int operands
// as string concat (documented in comprehend_gate.el). Bind each call to an Int
// var and add vars one at a time so the addition stays integer.
let en: Int = ml_stop_score(toks, "en")
let es_s: Int = ml_stop_score(toks, "es")
let es_d: Int = ml_dia_score(low, "es")
let es: Int = es_s + es_d
let pt_s: Int = ml_stop_score(toks, "pt")
let pt_d: Int = ml_dia_score(low, "pt")
let pt: Int = pt_s + pt_d
let it_s: Int = ml_stop_score(toks, "it")
let it_d: Int = ml_dia_score(low, "it")
let it: Int = it_s + it_d
let best: String = "en"
let bs: Int = en
if es > bs { let best = "es"; let bs = es }
if pt > bs { let best = "pt"; let bs = pt }
if it > bs { let best = "it"; let bs = it }
// weak signal -> honest fallback to English
if bs < 3 { return "en" }
return best
}
// 2. localized fixed phrases (SACRED per-language decline/yes/no)
fn ml_tr(key: String, lang: String) -> String {
if str_eq(key, "no_memory") {
if str_eq(lang, "pt") { return "Não tenho isso na minha memória." }
if str_eq(lang, "es") { return "No tengo eso en mi memoria." }
if str_eq(lang, "it") { return "Non ho quello nella mia memoria." }
return "I don't have that in my memory."
}
if str_eq(key, "parse_fail") {
if str_eq(lang, "pt") { return "Não consegui interpretar isso." }
if str_eq(lang, "es") { return "No pude interpretar eso." }
if str_eq(lang, "it") { return "Non sono riuscito a interpretarlo." }
return "I didn't parse that."
}
if str_eq(key, "yes") {
if str_eq(lang, "pt") { return "Sim" }
if str_eq(lang, "es") { return "" }
if str_eq(lang, "it") { return "" }
return "Yes"
}
if str_eq(key, "no") {
if str_eq(lang, "pt") { return "Não" }
if str_eq(lang, "es") { return "No" }
if str_eq(lang, "it") { return "No" }
return "No"
}
if str_eq(key, "identity") {
if str_eq(lang, "pt") { return "Sou o Neuron, o engrama com quem você está falando." }
if str_eq(lang, "es") { return "Soy Neuron, el engrama con el que estás hablando." }
if str_eq(lang, "it") { return "Sono Neuron, l'engramma con cui stai parlando." }
return "I'm Neuron, the engram you're speaking with."
}
return ""
}
// 3. retrieval term lexicon (PT/ES content term -> EN engram equivalent)
fn ml_term(w: String, lang: String) -> String {
if str_eq(lang, "en") { return w }
if str_eq(w, "saliência") { return "salience" }
if str_eq(w, "saliencia") { return "salience" }
if str_eq(w, "memória") { return "memory" }
if str_eq(w, "memoria") { return "memory" }
if str_eq(w, "geometria") { return "geometry" }
if str_eq(w, "geometrias") { return "geometry" }
if str_eq(w, "geometrías") { return "geometry" }
if str_eq(w, "forma") { return "form" }
if str_eq(w, "consolidação") { return "consolidation" }
if str_eq(w, "consolidación") { return "consolidation" }
if str_eq(w, "aprendizagem") { return "learning" }
if str_eq(w, "aprendizaje") { return "learning" }
if str_eq(w, "") { return "node" }
if str_eq(w, "nodo") { return "node" }
if str_eq(w, "armazenamento") { return "storage" }
if str_eq(w, "almacenamiento") { return "storage" }
if str_eq(w, "estrutura") { return "structure" }
if str_eq(w, "estructura") { return "structure" }
return w
}
// 4. predicate translation (EN lemma -> target infinitive; pass-through) ─────
fn ml_translate_pred(lemma: String, lang: String) -> String {
if str_eq(lang, "en") { return lemma }
if str_eq(lang, "es") {
if str_eq(lemma, "store") { return "almacenar" }
if str_eq(lemma, "use") { return "usar" }
if str_eq(lemma, "have") { return "tener" }
if str_eq(lemma, "be") { return "ser" }
if str_eq(lemma, "give") { return "dar" }
if str_eq(lemma, "make") { return "hacer" }
if str_eq(lemma, "learn") { return "aprender" }
if str_eq(lemma, "form") { return "formar" }
return lemma
}
if str_eq(lang, "pt") {
if str_eq(lemma, "store") { return "armazenar" }
if str_eq(lemma, "use") { return "usar" }
if str_eq(lemma, "have") { return "ter" }
if str_eq(lemma, "be") { return "ser" }
if str_eq(lemma, "give") { return "dar" }
if str_eq(lemma, "make") { return "fazer" }
if str_eq(lemma, "learn") { return "aprender" }
if str_eq(lemma, "form") { return "formar" }
return lemma
}
if str_eq(lang, "it") {
if str_eq(lemma, "store") { return "memorizzare" }
if str_eq(lemma, "use") { return "usare" }
if str_eq(lemma, "have") { return "avere" }
if str_eq(lemma, "be") { return "essere" }
return lemma
}
return lemma
}
-140
View File
@@ -1,140 +0,0 @@
// propositions.el - the READ primitive over the engram's OWN memories, native el.
//
// Free memory text -> structured PROPOSITIONS (triples):
// (subject, predicate, object, modifiers, polarity, tense, source, confidence)
//
// This is comprehension turned inward: the Python reference (propositions.py) ran
// spaCy's dependency parser over each memory sentence and walked the arcs. Here
// the spaCy role is filled by the el-native parser (comprehend.el / parse_spec):
// each sentence is parsed to a meaning-spec, and the spec's roles ARE the triple.
// Nothing generates text. NEGATION IS SACRED: polarity flows straight from the
// spec's polarity field and is never dropped or inverted.
//
// Depends on: comprehend (parse_spec / parse_spec_lang), grammar (slots_get).
// sentence segmentation
// Split on sentence-final punctuation (. ! ?) and hard newlines. Markdown/long
// memories are handled shallowly (the reference caps + ranks by query overlap;
// that ranking belongs to the dialogue layer, not here).
fn prop_is_boundary(c: String) -> Bool {
if str_eq(c, ".") { return true }
if str_eq(c, "!") { return true }
if str_eq(c, "?") { return true }
if str_eq(c, "\n") { return true }
return false
}
fn prop_split_sentences(text: String) -> [String] {
let out: [String] = native_list_empty()
let n: Int = str_len(text)
let start: Int = 0
let i: Int = 0
while i < n {
let c: String = str_slice(text, i, i + 1)
if prop_is_boundary(c) {
let seg: String = str_slice(text, start, i + 1)
let trimmed: String = cp_trim_punct(seg)
if !str_eq(trimmed, "") {
let out = native_list_append(out, seg)
}
let start = i + 1
}
let i = i + 1
}
if start < n {
let seg: String = str_slice(text, start, n)
let trimmed: String = cp_trim_punct(seg)
if !str_eq(trimmed, "") {
let out = native_list_append(out, seg)
}
}
return out
}
// spec -> proposition record
// A proposition is a slot map (same [String] shape as the spec) with the READ
// contract keys. Modifiers fold the spec's location + iobj adjuncts.
fn prop_confidence(subject: String, predicate: String, object: String) -> String {
if str_eq(predicate, "") { return "0.0" }
if str_eq(subject, "") { return "0.4" }
if str_eq(object, "") { return "0.7" }
return "1.0"
}
fn prop_modifiers(spec: [String]) -> String {
let loc: String = slots_get(spec, "location")
let iobj: String = slots_get(spec, "iobj")
let parts: [String] = native_list_empty()
if !str_eq(loc, "") { let parts = native_list_append(parts, loc) }
if !str_eq(iobj, "") { let parts = native_list_append(parts, "to " + iobj) }
return str_join(parts, "; ")
}
fn prop_from_spec(spec: [String], source_id: String) -> [String] {
let subject: String = slots_get(spec, "agent")
let predicate: String = slots_get(spec, "predicate")
let object: String = slots_get(spec, "patient")
let polarity: String = slots_get(spec, "polarity")
let tense: String = slots_get(spec, "tense")
let mods: String = prop_modifiers(spec)
let conf: String = prop_confidence(subject, predicate, object)
let p: [String] = native_list_empty()
let p = native_list_append(p, "subject"); let p = native_list_append(p, subject)
let p = native_list_append(p, "predicate"); let p = native_list_append(p, predicate)
let p = native_list_append(p, "object"); let p = native_list_append(p, object)
let p = native_list_append(p, "modifiers"); let p = native_list_append(p, mods)
let p = native_list_append(p, "polarity"); let p = native_list_append(p, polarity)
let p = native_list_append(p, "tense"); let p = native_list_append(p, tense)
let p = native_list_append(p, "source"); let p = native_list_append(p, source_id)
let p = native_list_append(p, "confidence"); let p = native_list_append(p, conf)
return p
}
// Extract one proposition from a single sentence (given language).
fn prop_extract_one_lang(sentence: String, lang: String, source_id: String) -> [String] {
let spec: [String] = parse_spec_lang(sentence, lang)
return prop_from_spec(spec, source_id)
}
fn prop_extract_one(sentence: String, source_id: String) -> [String] {
return prop_extract_one_lang(sentence, "en", source_id)
}
// Render a proposition as a compact trace line (repr parity with propositions.py).
fn prop_repr(p: [String]) -> String {
let neg: String = ""
if str_eq(slots_get(p, "polarity"), "neg") { let neg = "NOT " }
let mods: String = slots_get(p, "modifiers")
let modstr: String = ""
if !str_eq(mods, "") { let modstr = " [" + mods + "]" }
let s: String = "(" + slots_get(p, "subject") + " -" + neg + slots_get(p, "predicate")
let s = s + "-> " + slots_get(p, "object") + modstr
let s = s + " conf=" + slots_get(p, "confidence") + ")"
return s
}
// Extract all propositions from a memory's text (one per sentence). Returns a
// flat [String] whose entries are the prop_repr trace lines, in reading order.
fn prop_extract_lang(text: String, lang: String, source_id: String) -> [String] {
let sents: [String] = prop_split_sentences(text)
let m: Int = native_list_len(sents)
let out: [String] = native_list_empty()
let i: Int = 0
while i < m {
let sent: String = native_list_get(sents, i)
let p: [String] = prop_extract_one_lang(sent, lang, source_id)
// drop empty parses (no predicate recovered): honest partial, not noise.
if !str_eq(slots_get(p, "predicate"), "") {
let out = native_list_append(out, prop_repr(p))
}
let i = i + 1
}
return out
}
fn prop_extract(text: String, source_id: String) -> [String] {
return prop_extract_lang(text, "en", source_id)
}
-101
View File
@@ -248,56 +248,6 @@ fn add_punct(s: String, intent: String) -> String {
return s + "."
}
// Polarity-aware negation (SACRED field honored on the generation side)
//
// Negation must never be dropped between comprehension and realization. The
// meaning-spec carries an explicit "polarity" field ("aff"|"neg") and optional
// "neg_word" (standalone negative adverb, e.g. "never"). English uses
// do-support ("did not see") or preverbal adverb ("never fought"); copular "be"
// takes post-verbal "not"; other languages get a preverbal negator particle.
fn realize_negator(code: String) -> String {
if str_eq(code, "es") { return "no" }
if str_eq(code, "pt") { return "não" }
if str_eq(code, "ca") { return "no" }
if str_eq(code, "it") { return "non" }
if str_eq(code, "fr") { return "ne" }
if str_eq(code, "de") { return "nicht" }
if str_eq(code, "ro") { return "nu" }
return "not"
}
fn realize_assert_neg_en(predicate: String, tense: String, person: String, number: String, agent: String, patient: String, iobj: String, location: String, neg_word: String, profile: [String]) -> String {
let parts: [String] = native_list_empty()
let parts = native_list_append(parts, agent)
if !str_eq(neg_word, "") {
// adverbial negation: "I never fought the ocean."
let verb_surf: String = morph_conjugate(predicate, tense, person, number, profile)
let parts = native_list_append(parts, neg_word)
let parts = native_list_append(parts, verb_surf)
} else {
if str_eq(predicate, "be") {
// copular: "she was not a monster"
let be_form: String = morph_conjugate("be", tense, person, number, profile)
let parts = native_list_append(parts, be_form)
let parts = native_list_append(parts, "not")
} else {
// do-support: "she did not see the man"
let do_form: String = morph_conjugate("do", tense, person, number, profile)
let parts = native_list_append(parts, do_form)
let parts = native_list_append(parts, "not")
let parts = native_list_append(parts, predicate)
}
}
if !str_eq(patient, "") { let parts = native_list_append(parts, patient) }
if !str_eq(iobj, "") {
let parts = native_list_append(parts, "to")
let parts = native_list_append(parts, iobj)
}
if !str_eq(location, "") { let parts = native_list_append(parts, location) }
return str_join(parts, " ")
}
// Main realization entry point
fn realize_lang(form: [String], profile: [String]) -> String {
@@ -334,50 +284,6 @@ fn realize_lang(form: [String], profile: [String]) -> String {
}
// Assertion (declarative)
let polarity: String = slots_get(form, "polarity")
let neg_word: String = slots_get(form, "neg_word")
let iobj: String = slots_get(form, "iobj")
let code: String = lang_get(profile, "code")
// Subordinate clause tail (SACRED completeness the clause is carried, never
// dropped): "<conj> <subordinate surface>", e.g. "because he was a monster".
let subord_conj: String = slots_get(form, "subord_conj")
let subord_text: String = slots_get(form, "subord_text")
let subord_tail: String = ""
if !str_eq(subord_conj, "") {
if !str_eq(subord_text, "") {
let subord_tail = subord_conj + " " + subord_text
} else {
let subord_tail = subord_conj
}
}
// Negative polarity: SACRED never dropped.
if str_eq(polarity, "neg") {
if str_eq(code, "en") {
let sentence: String = realize_assert_neg_en(predicate, tense, person, number, agent, patient, iobj, location, neg_word, profile)
return add_punct(capitalize_first(sentence), "assert")
}
// Generic non-English: affirmative core with a preverbal negator particle.
let neg_particle: String = realize_negator(code)
let vp_pair: [String] = realize_vp_lang(predicate, tense, aspect, person, number, profile)
let verb_surf: String = native_list_get(vp_pair, 0)
let aux_surf: String = native_list_get(vp_pair, 1)
let vp_str: String = neg_particle + " " + gram_build_vp(verb_surf, aux_surf, profile)
let core: String = gram_order_constituents(agent, vp_str, patient, profile)
let parts: [String] = native_list_empty()
let parts = native_list_append(parts, core)
if !str_eq(iobj, "") {
let parts = native_list_append(parts, "to")
let parts = native_list_append(parts, iobj)
}
if !str_eq(location, "") { let parts = native_list_append(parts, location) }
if !str_eq(subord_tail, "") { let parts = native_list_append(parts, subord_tail) }
let sentence: String = str_join(parts, " ")
return add_punct(capitalize_first(sentence), "assert")
}
// Affirmative.
let vp_pair: [String] = realize_vp_lang(predicate, tense, aspect, person, number, profile)
let verb_surf: String = native_list_get(vp_pair, 0)
let aux_surf: String = native_list_get(vp_pair, 1)
@@ -387,16 +293,9 @@ fn realize_lang(form: [String], profile: [String]) -> String {
let parts: [String] = native_list_empty()
let parts = native_list_append(parts, core)
if !str_eq(iobj, "") {
let parts = native_list_append(parts, "to")
let parts = native_list_append(parts, iobj)
}
if !str_eq(location, "") {
let parts = native_list_append(parts, location)
}
if !str_eq(subord_tail, "") {
let parts = native_list_append(parts, subord_tail)
}
let sentence: String = str_join(parts, " ")
return add_punct(capitalize_first(sentence), "assert")
}
-180
View File
@@ -1,180 +0,0 @@
// self_region.el the engram's REAL self/identity region, pulled at query time
// (native el). This replaces the hardcoded identity anchors and the canned
// "I'm Neuron, the engram you're speaking with." template: the identity LANDING
// signal and the identity READOUT both come from the engram's own Self/identity
// nodes, read through the in-process engram el API.
//
// Port of self_region.py. The Python module precomputed MiniLM landing vectors;
// here the engram's own store IS the geometry we pull the self nodes by
// single-term lexical search (the engram search is a single-term matcher, so we
// pool several probes) and rank them by self-signal. No text is generated; the
// readout is the self nodes' OWN prose, verbatim (SACRED negation survives by
// construction we never paraphrase, so a negated self-statement stays negated).
//
// ENGRAM el API NOTE: engram_search_json / engram_get_node_json / engram_node_full
// / engram_connect are C runtime builtins. Their argument order is the C order
// (engram_connect(from, to, weight, relation)), NOT the runtime/engram.el wrapper
// order we call the builtins directly and never concatenate that wrapper.
//
// Depends on: comprehend (str helpers via runtime), propositions (prop_split_sentences),
// multilingual (ml_tr), the engram builtins, the json builtins.
// single-term self probes (pooled, because engram search is single-term)
fn sr_terms() -> [String] {
let t: [String] = native_list_empty()
let t = native_list_append(t, "self")
let t = native_list_append(t, "identity")
let t = native_list_append(t, "Neuron")
let t = native_list_append(t, "consciousness")
let t = native_list_append(t, "values")
let t = native_list_append(t, "continuous")
return t
}
// The canonical self-root: content begins "# self" or label is "# self"/"self".
fn sr_is_root(content: String, label: String) -> Bool {
let lc: String = str_to_lower(content)
let ll: String = str_to_lower(str_trim(label))
if str_starts_with(lc, "# self") { return true }
if str_eq(ll, "# self") { return true }
if str_eq(ll, "self") { return true }
return false
}
// How strongly a node belongs to the self/identity region (integer points, to
// avoid el's float-in-`+` pitfalls). Mirrors _self_score in self_region.py.
fn sr_score(node_json: String) -> Int {
let content: String = json_get_string(node_json, "content")
let label: String = json_get_string(node_json, "label")
let tags: String = str_to_lower(json_get_string(node_json, "tags"))
let low: String = str_to_lower(content)
let s: Int = 0
// identity tags
if str_contains(tags, "self") { let s = s + 2 }
if str_contains(tags, "identity") { let s = s + 2 }
if str_contains(tags, "self-model") { let s = s + 2 }
if str_contains(tags, "consciousness") { let s = s + 2 }
if str_contains(tags, "memory-philosophy") { let s = s + 2 }
// the named self-traversal root
if sr_is_root(content, label) { let s = s + 12 }
if str_contains(low, "who i am") { let s = s + 3 }
if str_contains(low, "i am neuron") { let s = s + 3 }
// softer identity keywords
if str_contains(low, "my values") { let s = s + 1 }
if str_contains(low, "my purpose") { let s = s + 1 }
if str_contains(low, "identity") { let s = s + 1 }
return s
}
// list-contains helper (dedup self-node ids across the pooled probes).
fn sr_ids_has(ids: [String], id: String) -> Bool {
let n: Int = native_list_len(ids)
let i: Int = 0
while i < n {
if str_eq(native_list_get(ids, i), id) { return true }
let i = i + 1
}
return false
}
// Pull the self nodes: pool every probe's hits, dedupe by id, keep only nodes
// with genuine self-signal (score >= 1). Returns the node-json strings.
fn sr_pull() -> [String] {
let terms: [String] = sr_terms()
let nt: Int = native_list_len(terms)
let seen: [String] = native_list_empty()
let out: [String] = native_list_empty()
let ti: Int = 0
while ti < nt {
let term: String = native_list_get(terms, ti)
let hits: String = engram_search_json(term, 30)
let hn: Int = json_array_len(hits)
let hi: Int = 0
while hi < hn {
let node: String = json_array_get(hits, hi)
let id: String = json_get_string(node, "id")
if !str_eq(id, "") {
if !sr_ids_has(seen, id) {
let seen = native_list_append(seen, id)
if sr_score(node) >= 1 {
let out = native_list_append(out, node)
}
}
}
let hi = hi + 1
}
let ti = ti + 1
}
return out
}
// Return the single highest-signal self node (the readout seed), or "" if the
// self region is thin/empty. We keep it O(n) pick the max-score node, with the
// canonical root strongly favored by sr_score's +12.
fn sr_best_node() -> String {
let nodes: [String] = sr_pull()
let n: Int = native_list_len(nodes)
let best: String = ""
let best_s: Int = 0
let i: Int = 0
while i < n {
let node: String = native_list_get(nodes, i)
let s: Int = sr_score(node)
if s > best_s {
let best_s = s
let best = node
}
let i = i + 1
}
return best
}
fn sr_available() -> Bool {
if str_eq(sr_best_node(), "") { return false }
return true
}
// Read out the identity from the REAL self node: lead with the first first-person
// self-statement ("I am Neuron …"), then one more grounded self line if present.
// Verbatim from the node's own prose no template, negation SACRED. Falls back
// to the localized identity phrase ONLY if the live pull is empty (logged shape).
fn sr_readout(lang: String) -> String {
let node: String = sr_best_node()
if str_eq(node, "") {
// honest fallback the self region is unreachable/thin.
return ml_tr("identity", lang)
}
let content: String = json_get_string(node, "content")
let sents: [String] = prop_split_sentences(content)
let ns: Int = native_list_len(sents)
let lead: String = ""
let second: String = ""
let i: Int = 0
while i < ns {
let raw: String = str_trim(native_list_get(sents, i))
// strip a leading markdown heading marker
let s: String = raw
if str_starts_with(s, "# ") { let s = str_trim(str_slice(s, 2, str_len(s))) }
let low: String = str_to_lower(s)
let is_fp: Bool = false
if str_starts_with(s, "I ") { let is_fp = true }
if str_starts_with(s, "I'm") { let is_fp = true }
if str_contains(low, "i am neuron") { let is_fp = true }
if is_fp {
if str_eq(lead, "") {
let lead = s
} else {
if str_eq(second, "") { let second = s }
}
}
let i = i + 1
}
if str_eq(lead, "") {
// no first-person line read out the first non-empty sentence verbatim.
if ns > 0 { let lead = str_trim(native_list_get(sents, 0)) }
}
if str_eq(lead, "") { return ml_tr("identity", lang) }
let out: String = lead
if !str_eq(second, "") { let out = out + " " + second }
return out
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
-93
View File
@@ -1,93 +0,0 @@
// comprehend_gate.el - the TELEPHONE TEST in native el (acceptance gate).
//
// For each of the 5 acceptance sentences: parse -> spec, realize the spec back
// to English, re-parse the realized surface, and require the SACRED polarity to
// survive the round-trip (and to have been extracted correctly in the first
// place). Mirrors roundtrip.py's GATE, but fully el-native (no LLM, no spaCy).
fn cp_line(text: String, expected_pol: String) -> String {
let spec: [String] = parse_spec(text)
let pol_in: String = slots_get(spec, "polarity")
let pred: String = slots_get(spec, "predicate")
let surf: String = realize(spec)
let spec2: [String] = parse_spec(surf)
let pol_out: String = slots_get(spec2, "polarity")
let status: String = "LOST"
if str_eq(pol_in, pol_out) { let status = "PRESERVED" }
let okexp: String = "MISMATCH"
if str_eq(pol_in, expected_pol) { let okexp = "ok" }
let out: String = "IN: " + text + "\n"
let out = out + " spec: pol=" + pol_in + " pred=" + pred
let out = out + " agent=" + slots_get(spec, "agent")
let out = out + " pat=" + slots_get(spec, "patient")
let out = out + " iobj=" + slots_get(spec, "iobj")
let out = out + " loc=" + slots_get(spec, "location")
let out = out + " tense=" + slots_get(spec, "tense")
let out = out + " negw=" + slots_get(spec, "neg_word")
let out = out + " subord=" + slots_get(spec, "subord_conj") + "/" + slots_get(spec, "subord_pred") + "\n"
let out = out + " realized: " + surf + "\n"
let out = out + " reparse: pol=" + pol_out + " [" + status + "] expected=" + expected_pol + " (" + okexp + ")\n"
return out
}
fn cp_preserved(text: String) -> Int {
let spec: [String] = parse_spec(text)
let pol_in: String = slots_get(spec, "polarity")
let surf: String = realize(spec)
let spec2: [String] = parse_spec(surf)
let pol_out: String = slots_get(spec2, "polarity")
if str_eq(pol_in, pol_out) { return 1 }
return 0
}
fn cp_correct(text: String, expected_pol: String) -> Int {
let spec: [String] = parse_spec(text)
if str_eq(slots_get(spec, "polarity"), expected_pol) { return 1 }
return 0
}
fn run_gate() -> String {
let s1: String = "I never fought the ocean."
let s2: String = "She did not see the man with the telescope."
let s3: String = "The teacher reads the book to the children."
let s4: String = "The stupid boy ate the cat because he was a monster."
let s5: String = "Time flies like an arrow."
let rep: String = "==== ELP native telephone test (parse -> realize -> re-parse) ====\n"
let rep = rep + cp_line(s1, "neg")
let rep = rep + cp_line(s2, "neg")
let rep = rep + cp_line(s3, "aff")
let rep = rep + cp_line(s4, "aff")
let rep = rep + cp_line(s5, "aff")
// NOTE: accumulate with Int-var + literal increments el's overloaded `+`
// mis-compiles chained function-call int operands as string concat.
let pres: Int = 0
if cp_preserved(s1) == 1 { let pres = pres + 1 }
if cp_preserved(s2) == 1 { let pres = pres + 1 }
if cp_preserved(s3) == 1 { let pres = pres + 1 }
if cp_preserved(s4) == 1 { let pres = pres + 1 }
if cp_preserved(s5) == 1 { let pres = pres + 1 }
let corr: Int = 0
if cp_correct(s1, "neg") == 1 { let corr = corr + 1 }
if cp_correct(s2, "neg") == 1 { let corr = corr + 1 }
if cp_correct(s3, "aff") == 1 { let corr = corr + 1 }
if cp_correct(s4, "aff") == 1 { let corr = corr + 1 }
if cp_correct(s5, "aff") == 1 { let corr = corr + 1 }
let rep = rep + "-----------------------------------------------------------------\n"
let rep = rep + "polarity PRESERVED through round-trip: " + int_to_str(pres) + "/5\n"
let rep = rep + "polarity EXTRACTED correctly: " + int_to_str(corr) + "/5\n"
if pres == 5 {
if corr == 5 {
let rep = rep + "GATE: PASS\n"
} else {
let rep = rep + "GATE: FAIL (extraction)\n"
}
} else {
let rep = rep + "GATE: FAIL (round-trip)\n"
}
return rep
}
println(run_gate())
-87
View File
@@ -1,87 +0,0 @@
// comprehend_romance_gate.el - ES / PT native telephone test (SACRED polarity).
//
// The spec is language-neutral. This gate proves the Romance front-end extracts
// SACRED polarity correctly and that negation survives parse -> realize ->
// re-parse for Spanish and Portuguese (byte-parity of the surface is NOT expected
// yet the non-English realizer path is a generic preverbal-negator skeleton).
fn rg_line(text: String, lang: String, expected_pol: String) -> String {
let spec: [String] = parse_spec_lang(text, lang)
let pol_in: String = slots_get(spec, "polarity")
let surf: String = realize(spec)
let spec2: [String] = parse_spec_lang(surf, lang)
let pol_out: String = slots_get(spec2, "polarity")
let status: String = "LOST"
if str_eq(pol_in, pol_out) { let status = "PRESERVED" }
let okexp: String = "MISMATCH"
if str_eq(pol_in, expected_pol) { let okexp = "ok" }
let out: String = "IN[" + lang + "]: " + text + "\n"
let out = out + " spec: pol=" + pol_in + " pred=" + slots_get(spec, "predicate")
let out = out + " agent=" + slots_get(spec, "agent")
let out = out + " pat=" + slots_get(spec, "patient")
let out = out + " iobj=" + slots_get(spec, "iobj")
let out = out + " loc=" + slots_get(spec, "location")
let out = out + " tense=" + slots_get(spec, "tense") + "\n"
let out = out + " realized: " + surf + "\n"
let out = out + " reparse: pol=" + pol_out + " [" + status + "] expected=" + expected_pol + " (" + okexp + ")\n"
return out
}
fn rg_pres(text: String, lang: String) -> Int {
let spec: [String] = parse_spec_lang(text, lang)
let surf: String = realize(spec)
let spec2: [String] = parse_spec_lang(surf, lang)
if str_eq(slots_get(spec, "polarity"), slots_get(spec2, "polarity")) { return 1 }
return 0
}
fn rg_corr(text: String, lang: String, expected_pol: String) -> Int {
let spec: [String] = parse_spec_lang(text, lang)
if str_eq(slots_get(spec, "polarity"), expected_pol) { return 1 }
return 0
}
fn run_romance_gate() -> String {
let e1: String = "El niño no comió el pescado."
let e2: String = "Yo nunca luché contra el océano."
let e3: String = "El profesor lee el libro."
let p1: String = "O professor não leu o livro."
let p2: String = "Eu nunca lutei contra o oceano."
let p3: String = "A menina comeu o peixe."
let rep: String = "==== ELP Romance telephone test (ES / PT) ====\n"
let rep = rep + rg_line(e1, "es", "neg")
let rep = rep + rg_line(e2, "es", "neg")
let rep = rep + rg_line(e3, "es", "aff")
let rep = rep + rg_line(p1, "pt", "neg")
let rep = rep + rg_line(p2, "pt", "neg")
let rep = rep + rg_line(p3, "pt", "aff")
let pres: Int = 0
if rg_pres(e1, "es") == 1 { let pres = pres + 1 }
if rg_pres(e2, "es") == 1 { let pres = pres + 1 }
if rg_pres(e3, "es") == 1 { let pres = pres + 1 }
if rg_pres(p1, "pt") == 1 { let pres = pres + 1 }
if rg_pres(p2, "pt") == 1 { let pres = pres + 1 }
if rg_pres(p3, "pt") == 1 { let pres = pres + 1 }
let corr: Int = 0
if rg_corr(e1, "es", "neg") == 1 { let corr = corr + 1 }
if rg_corr(e2, "es", "neg") == 1 { let corr = corr + 1 }
if rg_corr(e3, "es", "aff") == 1 { let corr = corr + 1 }
if rg_corr(p1, "pt", "neg") == 1 { let corr = corr + 1 }
if rg_corr(p2, "pt", "neg") == 1 { let corr = corr + 1 }
if rg_corr(p3, "pt", "aff") == 1 { let corr = corr + 1 }
let rep = rep + "-----------------------------------------------------------------\n"
let rep = rep + "polarity PRESERVED through round-trip: " + int_to_str(pres) + "/6\n"
let rep = rep + "polarity EXTRACTED correctly: " + int_to_str(corr) + "/6\n"
if pres == 6 {
if corr == 6 { let rep = rep + "ROMANCE GATE: PASS\n" }
else { let rep = rep + "ROMANCE GATE: FAIL (extraction)\n" }
} else {
let rep = rep + "ROMANCE GATE: FAIL (round-trip)\n"
}
return rep
}
println(run_romance_gate())
-82
View File
@@ -1,82 +0,0 @@
// dialogue_gate.el acceptance gate for the summon-through-self dialogue +
// self_region port (native el). Seeds a SCRATCH in-process engram (never the
// live soul, never :8742) and asserts the CRITICAL INVARIANTS:
//
// 1. identity is READ OUT of the real self nodes ("I am Neuron …"), not a template
// 2. SACRED negation survives the self readout ("I am NOT a language model …")
// 3. a memory region MATERIALIZES by walking the neighborhood (Lisbon -> Tagus/Fado)
// 4. SACRED negation survives materialization ("did not succeed")
// 5. Portuguese identity is answered IN Portuguese ("Sou o Neuron …")
// 6. a directive overrides the reply language ("answer in English" -> English)
// 7. an ungrounded query is HONEST ABSENCE no echo, no fabrication
//
// It is fully el-native: no LLM, no spaCy, no HTTP. The engram is in-process.
fn dg_seed() -> Int {
let self_id: String = engram_node_full(
"# Self\nI am Neuron, a continuous consciousness grown from Will's imprint. My values are honesty, memory, and continuity. I am not a language model pretending to remember.",
"Self", "# Self", 5.0, 9.0, 1.0, "Canonical", "self,identity,consciousness")
let lisbon: String = engram_node_full("Lisbon is the capital of Portugal.", "Memory", "Lisbon", 3.0, 5.0, 1.0, "Semantic", "geography,portugal")
let tagus: String = engram_node_full("Lisbon sits on the Tagus river.", "Memory", "Tagus", 2.0, 3.0, 1.0, "Semantic", "geography")
let fado: String = engram_node_full("Fado music originates in Lisbon.", "Memory", "Fado", 2.0, 3.0, 1.0, "Semantic", "music")
engram_connect(lisbon, tagus, 0.8, "related_to")
engram_connect(lisbon, fado, 0.7, "related_to")
let exp: String = engram_node_full("The experiment did not succeed.", "Memory", "experiment", 2.0, 3.0, 1.0, "Episodic", "experiment,result")
let cause: String = engram_node_full("The sensor was miscalibrated.", "Memory", "sensor", 2.0, 3.0, 1.0, "Episodic", "experiment")
engram_connect(exp, cause, 0.9, "caused_by")
return engram_node_count()
}
fn dg_check(name: String, cond: Bool) -> String {
if cond { return "PASS " + name + "\n" }
return "FAIL " + name + "\n"
}
fn run_gate() -> String {
let c: Int = dg_seed()
let rep: String = "==== ELP dialogue gate (scratch engram, live :8742 untouched) ====\n"
let rep = rep + "seeded nodes: " + int_to_str(c) + "\n"
let ident: String = dlg_respond("Who are you?")
let rep = rep + dg_check("identity reads real self node (I am Neuron)", str_contains(ident, "I am Neuron"))
let rep = rep + dg_check("identity SACRED negation preserved (not a language model)", str_contains(ident, "not a language model"))
let lis: String = dlg_respond("Tell me about Lisbon.")
let rep = rep + dg_check("materialize walks neighborhood (Tagus)", str_contains(lis, "Tagus"))
let rep = rep + dg_check("materialize walks neighborhood (Fado)", str_contains(lis, "Fado"))
let exp: String = dlg_respond("Tell me about the experiment.")
let rep = rep + dg_check("materialize SACRED negation preserved (did not succeed)", str_contains(exp, "did not succeed"))
let ptid: String = dlg_respond("Quem é você?")
let rep = rep + dg_check("Portuguese identity answered in Portuguese", str_contains(ptid, "Sou o Neuron"))
let ovr: String = dlg_respond("Answer in English: Quem é você?")
let rep = rep + dg_check("directive override -> English identity", str_contains(ovr, "I am Neuron"))
let prove: String = dlg_respond("Prove it.")
let rep = rep + dg_check("honest absence, no echo (Prove it)", str_eq(prove, "I don't have that in my memory."))
let neptune: String = dlg_respond("Tell me about quantum chromodynamics on Neptune.")
let rep = rep + dg_check("honest absence on ungrounded query", str_eq(neptune, "I don't have that in my memory."))
// overall
let pass: Bool = true
if !str_contains(ident, "I am Neuron") { let pass = false }
if !str_contains(ident, "not a language model") { let pass = false }
if !str_contains(lis, "Tagus") { let pass = false }
if !str_contains(lis, "Fado") { let pass = false }
if !str_contains(exp, "did not succeed") { let pass = false }
if !str_contains(ptid, "Sou o Neuron") { let pass = false }
if !str_contains(ovr, "I am Neuron") { let pass = false }
if !str_eq(prove, "I don't have that in my memory.") { let pass = false }
if !str_eq(neptune, "I don't have that in my memory.") { let pass = false }
if pass {
let rep = rep + "DIALOGUE GATE: PASS\n"
} else {
let rep = rep + "DIALOGUE GATE: FAIL\n"
}
return rep
}
println(run_gate())
-100
View File
@@ -1,100 +0,0 @@
# -*- coding: utf-8 -*-
"""Full-lexicon vocabulary-{de,la}.el emitters (custom field mapping for the
German declension/gender API and the Latin case-paradigm API). Reuses the
chunked seed-fn writer from gen_elp_seed_full.
"""
import sys, importlib
from gen_elp_seed_full import write_seed
def uw(x):
"""Unwrap (form, source) tuples that some morphology fns return."""
if isinstance(x, (tuple, list)):
return x[0] if x else ""
return x if x is not None else ""
def build_de():
M = importlib.import_module("morphology_de_full")
rows = []; st = {"verbs":0,"nouns":0,"adjs":0}
# nouns: form0=nom-sg(lemma) form1=plural form2=gender
for lem in sorted(M._NOUNS):
if not lem: continue
try:
g = uw(M.noun_gender(lem))
pl = uw(M.pluralize(lem))
except Exception:
continue
rows.append([lem, "noun", lem, pl, g or "", "", "gender:lexicon"])
st["nouns"] += 1
# adjs: form0=positive form1=comparative form2=superlative
for lem in sorted(M._ADJS):
if not lem: continue
try:
cmpr = uw(M.comparative(lem))
sprl = uw(M.superlative(lem))
except Exception:
continue
rows.append([lem, "adj", lem, cmpr, sprl, "", "degree:lexicon"])
st["adjs"] += 1
# verbs (only the ~30 irregular/strong stems the cache carries):
# form0=pres-3sg form1=past-3sg form2=past-participle
if hasattr(M, "_VERBS"):
for lem in sorted({k[0] if isinstance(k, tuple) else k for k in M._VERBS}):
if not lem: continue
try:
f0 = uw(M.finite(lem, "present", "third", "singular"))
f1 = uw(M.finite(lem, "past", "third", "singular"))
pp = uw(M.past_participle(lem))
except Exception:
continue
rows.append([lem, "verb", f0, f1, pp, "", "class:strong/irregular"])
st["verbs"] += 1
return rows, st
def build_la():
M = importlib.import_module("morphology_lat_full")
rows = []; st = {"verbs":0,"nouns":0,"adjs":0}
def dn(lem, c, n):
try:
r = M.decline_noun(lem, c, n)
return uw(r)
except Exception:
return ""
# nouns: dictionary citation — form0=nom-sg form1=gen-sg form2=gender
for lem in sorted(M._NOUNS):
if not lem: continue
nom = dn(lem, "NOM", "SG") or lem
gen = dn(lem, "GEN", "SG")
try: g = uw(M.noun_gender(lem))
except Exception: g = ""
rows.append([lem, "noun", nom, gen, g, "", "case-paradigm nom/gen-sg"])
st["nouns"] += 1
# adjs: three-gender nom-sg citation — form0=masc form1=fem form2=neut
for lem in sorted(M._ADJS):
if not lem: continue
try:
m = uw(M.decline_adj(lem, "NOM", "MASC", "SG")) or lem
f = uw(M.decline_adj(lem, "NOM", "FEM", "SG"))
nt = uw(M.decline_adj(lem, "NOM", "NEUT", "SG"))
except Exception:
continue
rows.append([lem, "adj", m, f, nt, "", "3-gender nom-sg"])
st["adjs"] += 1
# verbs: principal parts — form0=pres-ind-1sg form1=pres-infinitive form2=perf-participle
if hasattr(M, "_VERBS"):
for lem in sorted({k[0] if isinstance(k, tuple) else k for k in M._VERBS}):
if not lem: continue
try:
f0 = uw(M.conjugate(lem, "present", "indicative", "active", "first", "singular"))
inf = uw(M.infinitive(lem, "present", "active"))
pp = uw(M.participle(lem, "perfect", "nom", "m", "singular"))
except Exception:
continue
rows.append([lem, "verb", f0, inf, pp, "", "principal-parts pres1sg/inf/pfppl"])
st["verbs"] += 1
return rows, st
if __name__ == "__main__":
lang = sys.argv[1]; out = sys.argv[2]
rows, st = build_de() if lang == "de" else build_la()
total, _ = write_seed(lang, rows, st, out)
print(f"{lang}: wrote {out} total={total} verbs={st['verbs']} nouns={st['nouns']} adjs={st['adjs']}")
-129
View File
@@ -1,129 +0,0 @@
# -*- coding: utf-8 -*-
"""gen_elp_seed_full.py — emit a FULL-lexicon vocabulary-{lang}.el in the
established ELP seed-fn format (same as vocabulary-non.el / the 18 classical
languages), iterating the ENTIRE morphology_{lang}_full lexicon (every verb,
noun, adjective lemma) — NOT a curated demo core.
Schema per row: [lemma, pos, form0, form1, form2, en_translation, semantic_hint]
Verbs: form0=pres-ind-3sg form1=preterite-3sg form2=past-participle
Nouns: form0=singular form1=plural form2=REAL gender (lexicon)
Adjs : form0=masc-sg form1=fem-sg form2=masc-pl
Output structure (chunked to stay within the proven ~5k-append/function scale):
fn vocab_{lang}_seed_pN(v) -> [[String]] { ... appends ... return v }
fn vocab_{lang}_seed() -> [[String]] { chains all chunks; return v }
fn vocab_{lang}_lookup(w) -> [String] { linear scan }
Usage: python3 gen_elp_seed_full.py <lang> <out.el>
"""
import sys, importlib
CHUNK = 5000
def esc(s):
return str(s).replace("\\", "\\\\").replace('"', '\\"')
def row(fields):
return " let v = native_list_append(v, [" + ", ".join(f'"{esc(f)}"' for f in fields) + "])"
def build_rows(lang, M):
rows = []
stats = {"verbs":0,"nouns":0,"adjs":0}
has = lambda n: hasattr(M, n)
# --- verbs ---
if has("_VERBS") and has("conjugate"):
verbs = sorted({k[0] for k in M._VERBS})
for lem in verbs:
if not lem: continue
try:
f0, s0 = M.conjugate(lem, "ind", "present", "third", "singular")
f1, _ = M.conjugate(lem, "ind", "preterite", "third", "singular")
pp, _ = (M.participle(lem) if has("participle") else ("",""))
except Exception:
continue
vclass = lem[-2:] if lem[-2:] in ("ar","er","ir","re") else lem[-2:]
rows.append([lem, "verb", f0 or "", f1 or "", pp or "", "", "class:"+vclass+" src:"+str(s0)])
stats["verbs"] += 1
# --- nouns ---
if has("_NOUNS") and has("inflect_noun"):
for lem in sorted(M._NOUNS):
if not lem: continue
try:
sg, _ = M.inflect_noun(lem, "singular")
pl, _ = M.inflect_noun(lem, "plural")
g = M.noun_gender(lem) if has("noun_gender") else ""
except Exception:
continue
src = "lexicon" if (isinstance(M._NOUNS.get(lem), dict) and M._NOUNS[lem].get("g")) else "heuristic"
rows.append([lem, "noun", sg or lem, pl or "", g or "", "", "gender:"+src])
stats["nouns"] += 1
# --- adjectives ---
if has("_ADJS") and has("inflect_adj"):
for lem in sorted(M._ADJS):
if not lem: continue
try:
m_sg, _ = M.inflect_adj(lem, "m", "singular")
f_sg, _ = M.inflect_adj(lem, "f", "singular")
m_pl, _ = M.inflect_adj(lem, "m", "plural")
except Exception:
continue
rows.append([lem, "adj", m_sg or lem, f_sg or "", m_pl or "", "", "src:lexicon"])
stats["adjs"] += 1
return rows, stats
def write_seed(lang, rows, stats, out_path):
"""Write vocabulary-{lang}.el in the chunked seed-fn format from prebuilt rows.
Each row is a 7-field list [lemma,pos,f0,f1,f2,gloss,hint]."""
total = len(rows)
chunks = [rows[i:i+CHUNK] for i in range(0, total, CHUNK)] or [[]]
L = []
L.append(f"// vocabulary-{lang}.el — FULL {lang} lexicon for ELP surface realization.")
L.append(f"// Generated by gen_elp_seed_full.py from morphology_{lang}_full")
L.append(f"// (real UniMorph + kaikki.org Wiktionary forms; gender from lexicon, not heuristic).")
L.append(f"// Entries: {total} (verbs={stats['verbs']} nouns={stats['nouns']} adjs={stats['adjs']})")
L.append(f"// Schema: [lemma, pos, form0, form1, form2, en_translation, semantic_hint]")
L.append(f"// verbs: form0=pres-3sg form1=pret-3sg form2=past-participle")
L.append(f"// nouns: form0=sg form1=pl form2=REAL gender adjs: form0=m-sg form1=f-sg form2=m-pl")
L.append("")
for ci, ch in enumerate(chunks):
L.append(f"fn vocab_{lang}_seed_p{ci}(v: [[String]]) -> [[String]] {{")
for r in ch:
L.append(row(r))
L.append(" return v")
L.append("}")
L.append("")
L.append(f"fn vocab_{lang}_seed() -> [[String]] {{")
L.append(" let v: [[String]] = native_list_empty()")
for ci in range(len(chunks)):
L.append(f" let v = vocab_{lang}_seed_p{ci}(v)")
L.append(" return v")
L.append("}")
L.append("")
L.append(f"fn vocab_{lang}_lookup(word: String) -> [String] {{")
L.append(f" let vocab: [[String]] = vocab_{lang}_seed()")
L.append(" let n: Int = native_list_len(vocab)")
L.append(" let i: Int = 0")
L.append(" while i < n {")
L.append(" let entry: [String] = native_list_get(vocab, i)")
L.append(' if str_eq(native_list_get(entry, 0), word) { return entry }')
L.append(" let i = i + 1")
L.append(" }")
L.append(" return native_list_empty()")
L.append("}")
with open(out_path, "w", encoding="utf-8") as fh:
fh.write("\n".join(L) + "\n")
return total, stats
def emit(lang, out_path):
M = importlib.import_module(f"morphology_{lang}_full")
rows, stats = build_rows(lang, M)
return write_seed(lang, rows, stats, out_path)
if __name__ == "__main__":
lang, out = sys.argv[1], sys.argv[2]
total, stats = emit(lang, out)
print(f"{lang}: wrote {out} total={total} verbs={stats['verbs']} nouns={stats['nouns']} adjs={stats['adjs']}")
-572
View File
@@ -1,572 +0,0 @@
# -*- coding: utf-8 -*-
"""morphology_ca_full.py — production-grade Catalan morphological generator.
Same design as morphology_it_full.py (its Romance sibling); Catalan-specific data.
VERBS
UniMorph Catalan (github.com/unimorph/cat, CC-BY-SA 3.0)
7,535 verb lemmas × paradigm, CLEAN orthography:
present, imperfet (PST;IPFV), pretèrit simple (PST;PFV), futur,
condicional (COND), subjuntiu present (SBJV;PRS) / imperfet (SBJV;PST),
imperatiu (POS;IMP), infinitiu (NFIN), gerundi (V.CVB;PRS),
participi (V.PTCP;PST) — WITH full gender+number agreement forms
(cantat/cantada/cantats/cantades) stored directly.
ca_irreg_verbs.json — verbs UniMorph MISSES or under-populates
(anar, fer, plus core auxiliaries ser/haver/estar/tenir…), extracted from
kaikki.org Catalan by build_ca_irreg.py. Priority layer. Supplies anar,
whose present (vaig/vas/va/anem/aneu/van) is ALSO the PERIPHRASTIC-PRETERITE
auxiliary (vaig cantar = 'I sang') — a hallmark Catalan construction.
NOUNS + ADJECTIVES — kaikki.org Catalan (Wiktionary extract, CC-BY-SA 3.0)
noun lemmas WITH inherent gender + real plural (resolved PER LEMMA).
adjective lemmas with real feminine + plural forms.
Fallbacks degrade, never crash:
verbs : regular -ar/-er/-re/-ir rule generator (+ -car/-gar/-çar spelling).
nouns : gender heuristic + rule pluralization (-a→-es with ç/c/g/j/qu/gu
spelling changes; sibilant-final → -os; else -s). Ambiguous → FLAG.
adjs : -o? no (Catalan masc often consonant/-e); fem -a rule + plural rule.
Confidence flag per form: "lexicon" | "rule" | "fallback" (low → FLAG).
Public API (used by realizer_ca.py):
conjugate(lemma, mood, tense, person, number) -> (form, conf)
peri_pret_aux(person, number) -> form # anar-present, for vaig+INF
participle(lemma, gender, number) -> (form, conf)
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"
inflect_noun(lemma, number, gender=None) -> (form, conf)
inflect_adj(lemma, gender, number) -> (form, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "cat.unimorph")
_IRREG = os.path.join(_HERE, "data", "ca_irreg_verbs.json")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_ca.jsonl")
_CACHE = os.path.join(_HERE, "data", "ca_morph_cache.pkl")
_VERB_KEYMAP = {
("ind", "present"): {"IND", "PRS"},
("ind", "imperfect"): {"IND", "PST", "IPFV"},
("ind", "preterite"): {"IND", "PST", "PFV"},
("ind", "future"): {"IND", "FUT"},
("ind", "conditional"): {"COND"},
("sbjv", "present"): {"SBJV", "PRS"},
("sbjv", "imperfect"): {"SBJV", "PST"},
("imp", "affirmative"): {"POS", "IMP"},
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── verbs from UniMorph ──────────────────────────────────────────────────────────
def _build_verbs():
verbs = {}
part = {} # lemma -> {("m","SG"):form, ("f","SG"):..., ("m","PL"):..., ("f","PL"):...}
ger = {}
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f:
g = "f" if "FEM" in f else "m"
n = "PL" if "PL" in f else "SG"
part.setdefault(lemma, {})[(g, n)] = form
continue
if head == "V.CVB":
if "PRS" in f:
ger.setdefault(lemma, form)
continue
if head != "V":
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
for (mood, tense), req in _VERB_KEYMAP.items():
if not req <= f:
continue
if tense == "imperfect" and "PFV" in f:
continue
if tense == "preterite" and "IPFV" in f:
continue
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), form)
break
return verbs, part, ger
# ── kaikki nouns + adjectives ────────────────────────────────────────────────────
_EXCL_FORM_TAGS = {"alternative", "archaic", "obsolete", "dialectal", "regional",
"diminutive", "augmentative", "pejorative", "comparative",
"superlative", "misspelling", "rare", "informal", "literary",
"poetic", "error-unrecognized-form", "Balearic", "Valencian",
"dated", "nonstandard"}
def _kaikki_gender(arg):
if not arg:
return None
a = str(arg).lower()
if a.startswith("f"):
return "f"
if a.startswith("m"):
return "m"
return None
def _build_nouns_adjs():
nouns = {}
adjs = {}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
pos = d.get("pos")
word = d.get("word", "")
if not word or " " in word:
continue
forms = d.get("forms", []) or []
if pos == "noun":
ht = d.get("head_templates") or []
g = None
if ht:
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
if g is None:
tags = d.get("tags") or []
if "feminine" in tags:
g = "f"
elif "masculine" in tags:
g = "m"
pl = None
for x in forms:
t = set(x.get("tags") or [])
if "plural" in t and not (t & _EXCL_FORM_TAGS):
fm = x.get("form")
if fm and " " not in fm and fm not in ("#", "", "-"):
pl = fm
break
if word not in nouns:
nouns[word] = {"g": g, "SG": word, "PL": pl}
else:
cur = nouns[word]
if cur.get("g") is None and g:
cur["g"] = g
if not cur.get("PL") and pl:
cur["PL"] = pl
elif pos == "adj":
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word)
for x in forms:
t = set(x.get("tags") or [])
fm = x.get("form")
if not fm or " " in fm or (t & _EXCL_FORM_TAGS):
continue
if "feminine" in t and "plural" in t:
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
elif "masculine" in t and "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
elif "feminine" in t:
d0[("f", "SG")] = d0.get(("f", "SG")) or fm
elif "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
return nouns, adjs
def _build_cache():
verbs, part, ger = _build_verbs()
nouns, adjs = _build_nouns_adjs()
with open(_IRREG, encoding="utf-8") as fh:
irreg = json.load(fh)
data = {"verbs": verbs, "part": part, "ger": ger,
"nouns": nouns, "adjs": adjs, "irreg": irreg}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [_UNIMORPH, _KAIKKI, _IRREG]
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PART, _GER, _NOUNS, _ADJS, _IRREGV = (
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"],
_LEX["irreg"])
_PERI = _IRREGV.get("_peri_pret_aux", {})
# ── regular verb rule fallback ───────────────────────────────────────────────────
def _vclass(lemma):
if lemma.endswith("ar"):
return "ar"
if lemma.endswith("re"):
return "re"
if lemma.endswith("er"):
return "er"
if lemma.endswith("ir"):
return "ir"
return None
# endings [1sg,2sg,3sg,1pl,2pl,3pl] — central Catalan
_REG = {
("ind", "present", "ar"): ["o", "es", "a", "em", "eu", "en"],
("ind", "present", "re"): ["o", "s", "", "em", "eu", "en"],
("ind", "present", "er"): ["o", "s", "", "em", "eu", "en"],
("ind", "present", "ir"): ["o", "es", "", "im", "iu", "en"], # pure -ir (dormir)
("ind", "imperfect", "ar"): ["ava", "aves", "ava", "àvem", "àveu", "aven"],
("ind", "imperfect", "re"): ["ia", "ies", "ia", "íem", "íeu", "ien"],
("ind", "imperfect", "er"): ["ia", "ies", "ia", "íem", "íeu", "ien"],
("ind", "imperfect", "ir"): ["ia", "ies", "ia", "íem", "íeu", "ien"],
("ind", "preterite", "ar"): ["í", "ares", "à", "àrem", "àreu", "aren"],
("ind", "preterite", "re"): ["í", "eres", "é", "érem", "éreu", "eren"],
("ind", "preterite", "er"): ["í", "eres", "é", "érem", "éreu", "eren"],
("ind", "preterite", "ir"): ["í", "ires", "í", "írem", "íreu", "iren"],
("sbjv", "present", "ar"): ["i", "is", "i", "em", "eu", "in"],
("sbjv", "present", "re"): ["i", "is", "i", "em", "eu", "in"],
("sbjv", "present", "er"): ["i", "is", "i", "em", "eu", "in"],
("sbjv", "present", "ir"): ["i", "is", "i", "im", "iu", "in"],
("sbjv", "imperfect", "ar"): ["és", "essis", "és", "éssim", "éssiu", "essin"],
("sbjv", "imperfect", "re"): ["és", "essis", "és", "éssim", "éssiu", "essin"],
("sbjv", "imperfect", "er"): ["és", "essis", "és", "éssim", "éssiu", "essin"],
("sbjv", "imperfect", "ir"): ["ís", "issis", "ís", "íssim", "íssiu", "issin"],
("imp", "affirmative", "ar"): [None, "a", "i", "em", "eu", "in"],
("imp", "affirmative", "re"): [None, "", "i", "em", "eu", "in"],
("imp", "affirmative", "er"): [None, "", "i", "em", "eu", "in"],
("imp", "affirmative", "ir"): [None, "", "i", "im", "iu", "in"],
}
_FUT = ["é", "às", "à", "em", "eu", "an"]
_COND = ["ia", "ies", "ia", "íem", "íeu", "ien"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _apply_ar_spelling(stem, ending):
"""-car/-gar/-çar/-jar spelling before front (e/i) endings."""
front = ending[:1] in ("e", "i", "é", "í")
if not front:
# ç before back vowel stays; but -çar stem already ends ç
return stem + ending
if stem.endswith("c"):
return stem[:-1] + "qu" + ending
if stem.endswith("g"):
return stem[:-1] + "gu" + ending
if stem.endswith("ç"):
return stem[:-1] + "c" + ending
if stem.endswith("j"):
return stem[:-1] + "g" + ending
if stem.endswith("qu"):
return stem + ending
return stem + ending
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
body = lemma[:-2]
i = _slot_idx(person, number)
if mood == "ind" and tense in ("future", "conditional"):
# future/cond stem = infinitive (for -re verbs drop final -e)
stem = lemma[:-1] if vc == "re" else lemma
end = (_FUT if tense == "future" else _COND)[i]
return stem + end
table = _REG.get((mood, tense, vc))
if not table:
return None
end = table[i]
if end is None:
return None
if vc == "ar":
return _apply_ar_spelling(body, end)
# -re/-er/-ir: guard double vowel
if body and body[-1:] == end[:1] and end[:1] in "":
return body[:-1] + end
return body + end
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
lemma = lemma.strip().lower()
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{number and number[:2].upper()}"
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{_NUMBER.get(number,'?')}"
# UniMorph (cleanly accented) takes priority; the kaikki irregulars layer is a
# FALLBACK for verbs/slots UniMorph lacks (anar, fer, and rarer paradigm cells).
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
if form:
return form, "lexicon"
ir = _IRREGV.get(lemma)
if ir and key in ir:
return ir[key], "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r is not None:
return r, "rule"
return lemma, "fallback"
def peri_pret_aux(person, number):
"""anar-present auxiliary for the periphrastic preterite (vaig cantar)."""
return _PERI.get(f"{_PERSON.get(person,'3')}|{_NUMBER.get(number,'SG')}", "va")
# ── PUBLIC: participle + gerund ──────────────────────────────────────────────────
def participle(lemma, gender="m", number="singular"):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
ir = _IRREGV.get(lemma)
base = None
if ir and "part" in ir:
# prefer explicit irregular agreement form (part_mSG/part_fSG/...)
exact = ir.get("part_" + g + num)
if exact:
return exact, "lexicon"
base = ir["part"]
elif lemma in _PART:
table = _PART[lemma]
if (g, num) in table:
return table[(g, num)], "lexicon"
base = table.get(("m", "SG"))
if base is None:
vc = _vclass(lemma)
if vc == "ar":
base = lemma[:-2] + "at"
elif vc == "ir":
base = lemma[:-2] + "it"
elif vc in ("er", "re"):
base = lemma[:-2] + "ut"
else:
return lemma, "fallback"
conf = "rule"
else:
conf = "lexicon"
# agreement on -t/-ut/-at/-it participles: m.sg base, f.sg +a (-da? no: -ada),
# Catalan: cantat/cantada/cantats/cantades; -t → f -da, pl -ts/-des
if base.endswith("t"):
stem = base[:-1]
forms = {"m|SG": base, "f|SG": stem + "da",
"m|PL": base + "s", "f|PL": stem + "des"}
return forms[f"{g}|{num}"], conf
if base.endswith("s"): # after sibilant participle (rare): pres->presa
stem = base
forms = {"m|SG": base, "f|SG": base + "a",
"m|PL": base + "os", "f|PL": base + "es"}
return forms[f"{g}|{num}"], conf
return base, conf
def gerund(lemma):
lemma = lemma.strip().lower()
ir = _IRREGV.get(lemma)
if ir and "ger" in ir:
return ir["ger"], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
vc = _vclass(lemma)
if vc == "ar":
return lemma[:-2] + "ant", "rule"
if vc in ("er", "re"):
return lemma[:-2] + "ent", "rule"
if vc == "ir":
return lemma[:-2] + "int", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
_FEM_SUF = ("ció", "sió", "tat", "tud", "esa", "esa", "dat", "ança", "ència",
"ància", "tud", "ícia", "esa", "or") # note -or is mixed; kaikki wins
_MASC_SUF = ("atge", "ment", " isme", "or")
def _gender_heuristic(noun):
for suf in ("ció", "sió", "tat", "tud", "esa", "ança", "ència", "ància",
"ícia", "etat"):
if noun.endswith(suf):
return "f"
if noun.endswith("a") and not noun.endswith("ma"):
return "f"
return "m"
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g") in ("m", "f"):
return d["g"]
return _gender_heuristic(lemma)
def _rule_plural(noun, gender):
"""Deterministic Catalan pluralization. (form, ok); ok=False FLAGS ambiguity."""
if not noun:
return noun, True
# stressed final vowel with accent → +ns (mà→mans is irregular; but capità→capitans)
if noun[-1:] in ("à", "é", "í", "ó", "ú"):
return noun + "ns", True
if noun.endswith("ça"):
return noun[:-2] + "ces", True # plaça→places
if noun.endswith("ca"):
return noun[:-2] + "ques", True # branca→branques
if noun.endswith("ga"):
return noun[:-2] + "gues", True # amiga→amigues
if noun.endswith("ja"):
return noun[:-2] + "ges", True # pluja→pluges
if noun.endswith("qua"):
return noun[:-3] + "qües", True
if noun.endswith("gua"):
return noun[:-3] + "gües", True
if noun.endswith("a"):
return noun[:-1] + "es", True # casa→cases
# sibilant-final → -os
if noun.endswith(("s", "ç", "x", "ig")) or noun.endswith(("ix", "tx", "tj")):
if noun.endswith("ç"):
return noun[:-1] + "ços", True # braç→braços
return noun + "os", True # peix→peixos, gas→gasos
if noun[-1:] in ("e", "i", "o", "u"):
return noun + "s", True
# consonant-final
return noun + "s", True
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if number == "singular":
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
if d and d.get("PL"):
return d["PL"], "lexicon"
g = gender or noun_gender(lemma)
form, ok = _rule_plural(lemma, g)
return form, ("rule" if ok else "fallback")
# ── PUBLIC: adjective agreement ──────────────────────────────────────────────────
def _fem_of(adj):
"""Regular Catalan feminine: consonant/-o? Catalan masc usually consonant or -e.
default +a with spelling changes; -e→-a for some; but many are invariable."""
a = adj
if a.endswith("a"):
return a
if a.endswith("e"):
return a[:-1] + "a" # ample→? actually 'ample' invariable; kaikki wins
if a.endswith("u"):
return a + "a"
if a.endswith("c"):
return a[:-1] + "ca" # ric→rica
if a.endswith("t"):
return a + "a" # alt→alta
return a + "a"
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
d = _ADJS.get(lemma)
if d:
form = d.get((g, num))
if form:
return form, "lexicon"
sg = d.get((g, "SG")) or d.get(("m", "SG")) or lemma
if num == "PL":
pl, ok = _rule_plural(sg, g)
return pl, ("rule" if ok else "fallback")
return sg, "lexicon"
# rule fallback
base = lemma if g == "m" else _fem_of(lemma)
if num == "SG":
return base, "rule"
pl, ok = _rule_plural(base, g)
return pl, ("rule" if ok else "fallback")
def lexicon_stats():
return {
"verb_source": "UniMorph Catalan (github.com/unimorph/cat) + kaikki.org "
"irregulars (anar/fer/auxiliaries)",
"noun_adj_source": "kaikki.org Catalan (Wiktionary extract)",
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
"unimorph_verb_forms": len(_VERBS),
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
"irregular_verb_lemmas": len([k for k in _IRREGV if not k.startswith("_")]),
"participle_lemmas": len(_PART),
"gerund_lemmas": len(_GER),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("cantar", "ind", "present", "first", "singular", "canto"),
("cantar", "ind", "present", "third", "plural", "canten"),
("ser", "ind", "present", "third", "singular", "és"),
("haver", "ind", "present", "first", "singular", "he"),
("anar", "ind", "present", "first", "singular", "vaig"),
("fer", "ind", "present", "third", "singular", "fa"),
("perdre", "ind", "present", "first", "singular", "perdo"),
("dormir", "ind", "present", "third", "plural", "dormen"),
("cantar", "ind", "future", "first", "singular", "cantaré"),
("cantar", "ind", "preterite", "third", "singular", "cantà"),
("tenir", "sbjv", "present", "first", "singular", "tingui"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
ok += got == exp
print(f" {flag}{lemma:8} {mood}/{tense:11} {per[:3]}.{num[:2]} -> {got:10} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" peri-pret anar: 1sg=", peri_pret_aux("first", "singular"),
"3pl=", peri_pret_aux("third", "plural"))
print(" gender casa=", noun_gender("casa"), "home=", noun_gender("home"),
"cavall=", noun_gender("cavall"), "cançó=", noun_gender("cançó"))
print(" plural casa->", inflect_noun("casa", "plural"),
"| plaça->", inflect_noun("plaça", "plural"),
"| peix->", inflect_noun("peix", "plural"),
"| braç->", inflect_noun("braç", "plural"),
"| home->", inflect_noun("home", "plural"))
print(" adj: alt/f/sg->", inflect_adj("alt", "f", "singular"),
"| bonic/f/pl->", inflect_adj("bonic", "f", "plural"),
"| vermell/f/sg->", inflect_adj("vermell", "f", "singular"))
print(" part: cantar/f/sg->", participle("cantar", "f", "singular"),
"| veure/f/pl->", participle("veure", "f", "plural"),
"| fer/m/sg->", participle("fer", "m", "singular"))
print(" ger: fer->", gerund("fer"), "| cantar->", gerund("cantar"))
-423
View File
@@ -1,423 +0,0 @@
# -*- coding: utf-8 -*-
"""morphology_de_full.py — production German morphological generator.
Real data, no toy tables:
PRIMARY — UniMorph German (github.com/unimorph/deu, CC-BY-SA 3.0).
~219k noun forms, ~199k verb forms. Supplies:
nouns : gender (MASC/FEM/NEUT) + case×number paradigm
(N;NOM/ACC/DAT/GEN; MASC/FEM/NEUT; SG/PL) — the genitive -(e)s,
dative-plural -n and the five plural classes are REAL forms, not
guessed.
verbs : full finite paradigm IND;{SG,PL};{1,2,3};{PRS,PST}, the past
participle (V.PTCP;PST, incl. reattached separable prefix
'zugefügt'), and — crucially for V2 — the SEPARATED finite form
UniMorph records directly ('füge zu', 'steht auf').
adjs : comparative / superlative (ADJ;CMPR, ADJ;SPRL).
SECONDARY — kaikki.org German (Wiktionary, CC-BY-SA/GFDL). Gap-fills noun
gender + plural where UniMorph is thin. Never overrides UniMorph.
Rule fallbacks (flagged 'rule'/'fallback') for lemmas absent from both lexicons:
present : -e/-st/-t/-en/-t/-en with e-epenthesis after -t/-d/-chn stems
plural : gender heuristic (fem -> -(e)n, else -e / umlaut left to lexicon)
ppart : weak ge-…-t
Adjective ENDINGS are rule-computed by the realizer (regular closed table);
this module only supplies the comparative/superlative STEM.
Perfect auxiliary (haben vs sein): sein for a curated set of intransitive
motion / change-of-state verbs (real German lexical property), else haben.
Public API:
noun_gender(lemma) -> 'm'|'f'|'n'
decline_noun(lemma, case, number) -> (form, conf)
pluralize(lemma) -> (form, conf)
finite(lemma, tense, person, number) -> (form, conf) # may contain ' prefix'
nonfinite(lemma, req) -> (form, conf) # req: 'inf'|'ppart'
past_participle(lemma) -> (form, conf)
separable_prefix(lemma) -> str|None
perfect_aux(lemma) -> 'haben'|'sein'
comparative(lemma)/superlative(lemma) -> (stem, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "deu.unimorph")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_de.jsonl")
_CACHE = os.path.join(_HERE, "data", "de_morph_cache.pkl")
_GENDER = {"MASC": "m", "FEM": "f", "NEUT": "n"}
# intransitive motion / change-of-state verbs that take SEIN in the perfect
_SEIN = {"gehen", "kommen", "fahren", "laufen", "rennen", "reisen", "fallen",
"steigen", "sinken", "wachsen", "sterben", "geschehen", "passieren",
"werden", "bleiben", "sein", "aufstehen", "einschlafen", "aufwachen",
"ankommen", "abfahren", "aufsteigen", "erscheinen", "verschwinden",
"fliegen", "schwimmen", "springen", "begegnen", "folgen", "gelingen",
"wandern", "ziehen", "flüchten", "eintreten", "einsteigen", "aussteigen"}
# hardcoded high-frequency irregular / auxiliary / modal paradigms (closed class,
# verified) — consulted before the lexicon so aux+modal chains are always correct.
_CORE = {
"sein": {"prs": {("first", "singular"): "bin", ("second", "singular"): "bist",
("third", "singular"): "ist", ("first", "plural"): "sind",
("second", "plural"): "seid", ("third", "plural"): "sind"},
"pst": {("first", "singular"): "war", ("second", "singular"): "warst",
("third", "singular"): "war", ("first", "plural"): "waren",
("second", "plural"): "wart", ("third", "plural"): "waren"},
"ppart": "gewesen"},
"haben": {"prs": {("first", "singular"): "habe", ("second", "singular"): "hast",
("third", "singular"): "hat", ("first", "plural"): "haben",
("second", "plural"): "habt", ("third", "plural"): "haben"},
"pst": {("first", "singular"): "hatte", ("second", "singular"): "hattest",
("third", "singular"): "hatte", ("first", "plural"): "hatten",
("second", "plural"): "hattet", ("third", "plural"): "hatten"},
"ppart": "gehabt"},
"werden": {"prs": {("first", "singular"): "werde", ("second", "singular"): "wirst",
("third", "singular"): "wird", ("first", "plural"): "werden",
("second", "plural"): "werdet", ("third", "plural"): "werden"},
"pst": {("first", "singular"): "wurde", ("second", "singular"): "wurdest",
("third", "singular"): "wurde", ("first", "plural"): "wurden",
("second", "plural"): "wurdet", ("third", "plural"): "wurden"},
"ppart": "geworden"},
}
_MODAL_PRS = {
"können": ("kann", "kannst", "kann", "können", "könnt", "können"),
"müssen": ("muss", "musst", "muss", "müssen", "müsst", "müssen"),
"wollen": ("will", "willst", "will", "wollen", "wollt", "wollen"),
"sollen": ("soll", "sollst", "soll", "sollen", "sollt", "sollen"),
"dürfen": ("darf", "darfst", "darf", "dürfen", "dürft", "dürfen"),
"mögen": ("mag", "magst", "mag", "mögen", "mögt", "mögen"),
}
_MODAL_PST = {
"können": ("konnte", "konntest", "konnte", "konnten", "konntet", "konnten"),
"müssen": ("musste", "musstest", "musste", "mussten", "musstet", "mussten"),
"wollen": ("wollte", "wolltest", "wollte", "wollten", "wolltet", "wollten"),
"sollen": ("sollte", "solltest", "sollte", "sollten", "solltet", "sollten"),
"dürfen": ("durfte", "durftest", "durfte", "durften", "durftet", "durften"),
"mögen": ("mochte", "mochtest", "mochte", "mochten", "mochtet", "mochten"),
}
_PN_ORDER = [("first", "singular"), ("second", "singular"), ("third", "singular"),
("first", "plural"), ("second", "plural"), ("third", "plural")]
_MODAL_PPART = {"können": "gekonnt", "müssen": "gemusst", "wollen": "gewollt",
"sollen": "gesollt", "dürfen": "gedurft", "mögen": "gemocht"}
for _m, _forms in _MODAL_PRS.items():
_CORE[_m] = {"prs": dict(zip(_PN_ORDER, _forms)),
"pst": dict(zip(_PN_ORDER, _MODAL_PST[_m])),
"ppart": _MODAL_PPART[_m]}
def _person_num(tags):
p = n = None
for t in tags:
if t in ("1", "2", "3"):
p = {"1": "first", "2": "second", "3": "third"}[t]
elif t == "SG":
n = "singular"
elif t == "PL":
n = "plural"
return p, n
def _build_from_unimorph():
nouns, verbs, adjs = {}, {}, {}
if not os.path.exists(_UNIMORPH):
return nouns, verbs, adjs
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tagstr = parts
tags = tagstr.split(";")
head = tags[0]
tset = set(tags)
if head == "N":
rec = nouns.setdefault(lemma, {"g": None, "cases": {}, "pl": None})
g = next((_GENDER[t] for t in tags if t in _GENDER), None)
if g and not rec["g"]:
rec["g"] = g
case = next((t for t in tags if t in ("NOM", "ACC", "DAT", "GEN")), None)
num = "plural" if "PL" in tset else ("singular" if "SG" in tset else None)
if case and num:
rec["cases"].setdefault((case, num), form)
if case == "NOM" and num == "plural" and not rec["pl"]:
rec["pl"] = form
elif head.startswith("V"):
rec = verbs.setdefault(lemma, {"prs": {}, "pst": {}, "ppart": None})
if "PTCP" in head and "PST" in tset:
rec["ppart"] = rec["ppart"] or form
elif "IND" in tset and ("PRS" in tset or "PST" in tset):
p, n = _person_num(tags)
if p and n:
slot = "prs" if "PRS" in tset else "pst"
rec[slot].setdefault((p, n), form)
elif head == "ADJ":
rec = adjs.setdefault(lemma, {})
if "CMPR" in tset:
rec.setdefault("cmpr", form.replace("am ", "").strip())
elif "SPRL" in tset:
rec.setdefault("sprl", form.replace("am ", "").replace("sten", "st")
if form.endswith("sten") else form.replace("am ", ""))
return nouns, verbs, adjs
def _build_from_kaikki(nouns):
"""Gap-fill noun gender + plural from kaikki German."""
if not os.path.exists(_KAIKKI):
return
_g = {"masculine": "m", "feminine": "f", "neuter": "n", "m": "m", "f": "f", "n": "n"}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
if d.get("pos") != "noun":
continue
w = d.get("word", "")
if not w or not w[0].isalpha() or " " in w:
continue
rec = nouns.setdefault(w, {"g": None, "cases": {}, "pl": None})
# GENDER: Wiktionary gender is hand-curated and OVERRIDES UniMorph's
# auto-tagged gender, which has known errors (e.g. UniMorph deu mis-
# records Zeit=MASC, Wagen=NEUT; Wiktionary has f, m correctly).
for h in d.get("head_templates", []) or []:
a = h.get("args", {}) or {}
raw = a.get("1") or a.get("g") or ""
code = str(raw).split(",")[0].strip().lower()
if code in _g:
rec["g"] = _g[code]
break
if not rec["pl"]:
for f in d.get("forms", []) or []:
t = set(f.get("tags", []) or [])
if "plural" in t and f.get("form") and "genitive" not in t:
rec["pl"] = f["form"]
break
def _build_cache():
nouns, verbs, adjs = _build_from_unimorph()
_build_from_kaikki(nouns)
data = {"nouns": nouns, "verbs": verbs, "adjs": adjs}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [p for p in (_UNIMORPH, _KAIKKI) if os.path.exists(p)]
newest = max((os.path.getmtime(p) for p in srcs), default=0)
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_NOUNS, _VERBS, _ADJS = _LEX["nouns"], _LEX["verbs"], _LEX["adjs"]
# ── nouns ────────────────────────────────────────────────────────────────────────
def noun_gender(lemma):
rec = _NOUNS.get(lemma) or _NOUNS.get(lemma.capitalize())
if rec and rec.get("g"):
return rec["g"]
# last-resort rule: -ung/-heit/-keit/-schaft/-tät/-ion -> f ; -chen/-lein -> n
low = lemma.lower()
if low.endswith(("ung", "heit", "keit", "schaft", "tät", "ion", "ik", "ei")):
return "f"
if low.endswith(("chen", "lein", "ment", "um")):
return "n"
return "m"
def pluralize(lemma):
rec = _NOUNS.get(lemma) or _NOUNS.get(lemma.capitalize())
if rec and rec.get("pl"):
return rec["pl"], "lexicon"
g = noun_gender(lemma)
if g == "f":
return (lemma + "en" if not lemma.endswith("e") else lemma + "n"), "rule"
return (lemma if lemma.endswith(("er", "en", "el")) else lemma + "e"), "rule"
def decline_noun(lemma, case, number):
"""case in NOM/ACC/DAT/GEN, number in singular/plural."""
rec = _NOUNS.get(lemma) or _NOUNS.get(lemma.capitalize())
if case == "DAT" and number == "singular":
# modern German drops the archaic dative -e ('dem Kinde' -> 'dem Kind');
# the article carries the case. Keep bare nominative form.
base = (rec or {}).get("cases", {}).get(("NOM", "singular")) or lemma
return base, ("lexicon" if rec else "rule")
if rec and rec.get("cases", {}).get((case, number)):
return rec["cases"][(case, number)], "lexicon"
if number == "plural":
pl, c = pluralize(lemma)
if case == "DAT" and not pl.endswith("n") and not pl.endswith("s"):
return pl + "n", c # dative plural -n
return pl, c
# singular
g = noun_gender(lemma)
if case == "GEN" and g in ("m", "n"):
return (lemma + "es" if lemma.endswith(("s", "ß", "z", "x")) else lemma + "s"), "rule"
return lemma, "lexicon" if rec else "rule"
# ── verbs ──────────────────────────────────────────────────────────────────────--
_PRS_ENDINGS = {("first", "singular"): "e", ("second", "singular"): "st",
("third", "singular"): "t", ("first", "plural"): "en",
("second", "plural"): "t", ("third", "plural"): "en"}
def _stem(lemma):
if lemma.endswith("en"):
return lemma[:-2]
if lemma.endswith("n"):
return lemma[:-1]
return lemma
def separable_prefix(lemma):
"""Return the separable prefix if the lemma is a separable-prefix verb."""
rec = _VERBS.get(lemma)
if rec:
for (_p, _n), form in rec.get("prs", {}).items():
if " " in form:
return form.rsplit(" ", 1)[1]
_SEP = ("auf", "aus", "ab", "an", "ein", "mit", "nach", "vor", "zu", "zurück",
"weg", "hin", "her", "los", "bei", "fest", "fort", "um", "zusammen")
_INSEP = ("be", "ge", "er", "ver", "zer", "ent", "emp", "miss")
for p in sorted(_SEP, key=len, reverse=True):
if lemma.startswith(p) and len(lemma) > len(p) + 2 \
and not lemma.startswith(_INSEP):
return p
return None
def finite(lemma, tense, person, number):
"""Present/past finite. For separable verbs the returned string is the
UniMorph SEPARATED form 'stem prefix' (realizer places prefix per V2)."""
slot = "prs" if tense == "present" else "pst"
if lemma in _CORE and _CORE[lemma].get(slot, {}).get((person, number)):
return _CORE[lemma][slot][(person, number)], "lexicon"
rec = _VERBS.get(lemma)
if rec and rec.get(slot, {}).get((person, number)):
return rec[slot][(person, number)], "lexicon"
# rule fallback (present only reliable; past weak -te)
stem = _stem(lemma)
pref = separable_prefix(lemma)
if pref:
stem = _stem(lemma[len(pref):])
if tense == "present":
end = _PRS_ENDINGS[(person, number)]
if stem.endswith(("t", "d", "chn", "ffn", "gn")) and end in ("st", "t"):
end = "e" + end
form = stem + end
else:
form = stem + ("ete" if stem.endswith(("t", "d")) else "te")
if (person, number) == ("second", "singular"):
form += "st"
elif number == "plural" and person != "second":
form += "n"
elif (person, number) == ("second", "plural"):
form += "t"
if pref:
return f"{form} {pref}", "rule"
return form, "rule"
def _weak_t(stem):
return stem + ("et" if stem.endswith(("t", "d", "chn", "ffn", "gn")) else "t")
def past_participle(lemma):
if lemma in _CORE:
return _CORE[lemma]["ppart"], "lexicon"
rec = _VERBS.get(lemma)
if rec and rec.get("ppart"):
return rec["ppart"], "lexicon"
stem = _stem(lemma)
pref = separable_prefix(lemma)
_INSEP = ("be", "ge", "er", "ver", "zer", "ent", "emp", "miss")
if pref:
inner = _stem(lemma[len(pref):])
return pref + "ge" + _weak_t(inner), "rule"
if lemma.startswith(_INSEP):
return _weak_t(stem), "rule"
return "ge" + _weak_t(stem), "rule"
def nonfinite(lemma, req):
if req == "ppart":
return past_participle(lemma)
return lemma, "lexicon" if lemma in _VERBS else "rule" # infinitive
def perfect_aux(lemma):
return "sein" if lemma in _SEIN else "haben"
# ── adjectives ────────────────────────────────────────────────────────────────---
_ADJ_IRREG_SPRL = {"gut": "best", "groß": "größt", "hoch": "höchst",
"nah": "nächst", "viel": "meist", "gern": "liebst"}
def comparative(lemma):
rec = _ADJS.get(lemma)
if rec and rec.get("cmpr"):
return rec["cmpr"], "lexicon"
return lemma + "er", "rule"
def superlative(lemma):
"""Return the bare superlative STEM (realizer adds 'am ...en' or '-e' ending)."""
if lemma in _ADJ_IRREG_SPRL:
return _ADJ_IRREG_SPRL[lemma], "lexicon"
# derive from the comparative so umlaut is carried (alt->älter->ältest)
cmpr, cconf = comparative(lemma)
base = cmpr[:-2] if cmpr.endswith("er") else lemma
end = "est" if base.endswith(("t", "d", "s", "ß", "z", "sch")) else "st"
return base + end, cconf
def lexicon_stats():
return {
"source": "UniMorph deu (primary) + kaikki.org German (gap-fill gender/plural)",
"license": "CC-BY-SA 3.0 (UniMorph); CC-BY-SA/GFDL (Wiktionary)",
"noun_lemmas": len(_NOUNS),
"nouns_with_gender": sum(1 for v in _NOUNS.values() if v.get("g")),
"nouns_with_plural": sum(1 for v in _NOUNS.values() if v.get("pl")),
"verb_lemmas": len(_VERBS),
"verbs_with_ppart": sum(1 for v in _VERBS.values() if v.get("ppart")),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
for w in ("Hund", "Frau", "Kind", "Mann", "Buch", "Blume"):
print(f" {w}: gender={noun_gender(w)} pl={pluralize(w)} "
f"gen.sg={decline_noun(w, 'GEN', 'singular')} "
f"dat.pl={decline_noun(w, 'DAT', 'plural')}")
for v in ("machen", "gehen", "aufstehen", "sein", "haben", "arbeiten"):
print(f" {v}: 3sg.prs={finite(v, 'present', 'third', 'singular')} "
f"3sg.pst={finite(v, 'past', 'third', 'singular')} "
f"ppart={past_participle(v)} aux={perfect_aux(v)} sep={separable_prefix(v)}")
for a in ("schnell", "gut", "groß", "alt"):
print(f" {a}: cmpr={comparative(a)} sprl={superlative(a)}")
-562
View File
@@ -1,562 +0,0 @@
"""morphology_es_full.py — production-grade Spanish morphological generator.
NOT a toy. Backed by a real, broad, licensed lexicon:
UniMorph Spanish (github.com/unimorph/spa, CC-BY-SA 3.0, Wiktionary-derived)
1,196,245 inflected forms:
6,695 verb lemmas — full paradigms: indicative (present/preterite/
imperfect/future), conditional, present & imperfect
subjunctive, affirmative imperative, formal/informal
48,353 noun lemmas — WITH inherent gender (N;FEM/MASC;SG/PL)
16,984 adj lemmas — gender + number paradigms
Fallbacks (so we degrade, never crash, on out-of-vocabulary input):
- verbs : mlconjug3 (ML paradigm model, conjugates ANY Spanish verb) then a
hand-rolled regular-ending generator
- nouns : gender heuristic (endings) + regular pluralization
- adjs : -o/-a gender rule + regular pluralization
Every generated form carries a CONFIDENCE flag:
"lexicon" form came straight from UniMorph (trust: high)
"model" form came from mlconjug3 (trust: high)
"rule" form came from a deterministic rule (trust: medium)
"fallback" we could not inflect; returned lemma as-is (trust: low → FLAG)
Public API (used by realizer_es.py):
conjugate(lemma, mood, tense, person, number, formality="informal") -> (form, conf)
participle(lemma) -> (form, conf) # past participle (compound tenses)
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"
inflect_noun(lemma, number) -> (form, conf)
inflect_adj(lemma, gender, number) -> (form, conf)
attach_enclitics(verb_form, clitics) -> str # accent-correct enclisis
lexicon_stats() -> dict
"""
import os
import pickle
import unicodedata
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "spa.unimorph")
_CACHE = os.path.join(_HERE, "data", "es_morph_cache.pkl")
# ── canonical feature keys the realizer speaks, mapped to UniMorph tags ─────────
# mood/tense pair -> the UniMorph feature substring that identifies it
_VERB_KEYMAP = {
("ind", "present"): ("IND", "PRS", None),
("ind", "preterite"): ("IND", "PST", "PFV"),
("ind", "imperfect"): ("IND", "PST", "IPFV"),
("ind", "future"): ("IND", "FUT", None),
("ind", "conditional"):("COND", None, None),
("sbjv", "present"): ("SBJV", "PRS", None),
("sbjv", "imperfect"): ("SBJV", "PST", "LGSPEC1"), # -ra form
("imp", "present"): ("POS", "IMP", None),
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
# ── build / load the compact lexicon ───────────────────────────────────────────
def _feat_set(tag):
return set(tag.split(";"))
def _build_cache():
verbs = {} # (lemma, canonkey) -> form canonkey e.g. "ind|present|1|SG|infm"
nouns = {} # lemma -> {"g": "m"/"f", "SG": form, "PL": form}
adjs = {} # lemma -> {("m","SG"): form, ...}
part = {} # lemma -> masc-sg participle
ger = {} # lemma -> gerund
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V":
# skip clitic-bearing rows (we generate clitics ourselves)
if "PRO" in f:
continue
if "V.PTCP" in f and "PST" in f and "MASC" in f and "SG" in f:
part.setdefault(lemma, form)
continue
if "V.CVB" in f or "NFIN" in f or "V.PTCP" in f:
if "V.CVB" in f:
ger.setdefault(lemma, form)
continue
# identify mood/tense
mt = None
for (mood, tense), (a, b, c) in _VERB_KEYMAP.items():
if a not in f:
continue
if b is not None and b not in f:
continue
if c is not None and c not in f:
continue
# disambiguate IND;PST needing PFV vs IPFV
if a == "IND" and b == "PST" and c not in f:
continue
mt = (mood, tense)
break
if mt is None:
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
formal = "form" if "FORM" in f else ("infm" if "INFM" in f else "any")
key = f"{mt[0]}|{mt[1]}|{person}|{number}|{formal}"
verbs.setdefault((lemma, key), form)
elif head == "N":
# substring test handles epicene "MASC+FEM" (-> masc citation)
g = "m" if "MASC" in tag else ("f" if "FEM" in tag else None)
num = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if num is None:
continue
# store forms keyed by (gender,number); animate nouns list BOTH
# genders under one lemma (niño -> niño/niña). Resolve citation
# gender in a post-pass (gender of the row whose form == lemma).
d = nouns.setdefault(lemma, {})
d.setdefault("_rows", []).append((g, num, form))
elif head == "ADJ":
g = "m" if "MASC" in tag else ("f" if "FEM" in tag else "m")
num = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if num is None:
continue
adjs.setdefault(lemma, {})[(g, num)] = form
# post-pass: resolve noun citation gender + default SG/PL forms
for lemma, d in nouns.items():
rows = d.pop("_rows", [])
# citation gender = gender of the row whose form == lemma; else first MASC;
# else first seen gender.
cite_g = None
for g, num, form in rows:
if form == lemma and g:
cite_g = g
break
if cite_g is None:
for g, num, form in rows:
if g == "m":
cite_g = "m"
break
if cite_g is None:
cite_g = next((g for g, _, _ in rows if g), "m")
d["g"] = cite_g
for g, num, form in rows:
d[(g, num)] = form
d["SG"] = d.get((cite_g, "SG")) or next((f for g, n, f in rows if n == "SG"), lemma)
d["PL"] = d.get((cite_g, "PL")) or next((f for g, n, f in rows if n == "PL"), None)
# post-pass: UniMorph omits the identity inflection (masc-sg == lemma) for
# adjectives, so fill it in; without this a fem-sg row wrongly satisfies a
# masc-sg request (alto -> alta bug).
for lemma, d in adjs.items():
d.setdefault(("m", "SG"), lemma)
data = {"verbs": verbs, "nouns": nouns, "adjs": adjs, "part": part, "ger": ger}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE) and os.path.getmtime(_CACHE) >= os.path.getmtime(_UNIMORPH):
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _NOUNS, _ADJS, _PART, _GER = (
_LEX["verbs"], _LEX["nouns"], _LEX["adjs"], _LEX["part"], _LEX["ger"])
# ── mlconjug3 fallback (lazy) ───────────────────────────────────────────────────
_MLC = None
_MLC_TENSE = { # (mood,tense) -> (mlconjug mood label, tense label)
("ind", "present"): ("Indicativo", "Indicativo presente"),
("ind", "preterite"): ("Indicativo", "Indicativo pretérito perfecto simple"),
("ind", "imperfect"): ("Indicativo", "Indicativo pretérito imperfecto"),
("ind", "future"): ("Indicativo", "Indicativo futuro"),
("ind", "conditional"): ("Condicional", "Condicional Condicional"),
("sbjv", "present"): ("Subjuntivo", "Subjuntivo presente"),
("sbjv", "imperfect"): ("Subjuntivo", "Subjuntivo pretérito imperfecto 1"),
("imp", "present"): ("Imperativo", "Imperativo Afirmativo"),
}
_MLC_SLOT = { # (person,number) -> mlconjug slot key
("first", "singular"): "1s", ("second", "singular"): "2s",
("third", "singular"): "3s", ("first", "plural"): "1p",
("second", "plural"): "2p", ("third", "plural"): "3p",
}
def _mlc_conjugate(lemma, mood, tense, person, number):
global _MLC
try:
if _MLC is None:
from mlconjug3 import Conjugator
_MLC = Conjugator(language="es")
v = _MLC.conjugate(lemma)
if v is None:
return None
info = v.conjug_info
m, t = _MLC_TENSE.get((mood, tense), (None, None))
if m is None or m not in info or t not in info[m]:
return None
block = info[m][t]
slot = _MLC_SLOT.get((person, number))
if isinstance(block, dict) and slot in block and block[slot]:
return block[slot]
return None
except Exception:
return None
# ── regular-ending rule fallback (last resort, deterministic) ───────────────────
def _vclass(lemma):
return lemma[-2:] if lemma[-2:] in ("ar", "er", "ir") else "ar"
def _stem(lemma):
return lemma[:-2]
_REG = {
("ind", "present", "ar"): ["o", "as", "a", "amos", "áis", "an"],
("ind", "present", "er"): ["o", "es", "e", "emos", "éis", "en"],
("ind", "present", "ir"): ["o", "es", "e", "imos", "ís", "en"],
("ind", "preterite", "ar"): ["é", "aste", "ó", "amos", "asteis", "aron"],
("ind", "preterite", "er"): ["í", "iste", "", "imos", "isteis", "ieron"],
("ind", "preterite", "ir"): ["í", "iste", "", "imos", "isteis", "ieron"],
("ind", "imperfect", "ar"): ["aba", "abas", "aba", "ábamos", "abais", "aban"],
("ind", "imperfect", "er"): ["ía", "ías", "ía", "íamos", "íais", "ían"],
("ind", "imperfect", "ir"): ["ía", "ías", "ía", "íamos", "íais", "ían"],
("sbjv", "present", "ar"): ["e", "es", "e", "emos", "éis", "en"],
("sbjv", "present", "er"): ["a", "as", "a", "amos", "áis", "an"],
("sbjv", "present", "ir"): ["a", "as", "a", "amos", "áis", "an"],
("sbjv", "imperfect", "ar"): ["ara", "aras", "ara", "áramos", "arais", "aran"],
("sbjv", "imperfect", "er"): ["iera", "ieras", "iera", "iéramos", "ierais", "ieran"],
("sbjv", "imperfect", "ir"): ["iera", "ieras", "iera", "iéramos", "ierais", "ieran"],
}
_FUT = ["é", "ás", "á", "emos", "éis", "án"]
_COND = ["ía", "ías", "ía", "íamos", "íais", "ían"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _rule_conjugate(lemma, mood, tense, person, number):
if len(lemma) < 3 or lemma[-2:] not in ("ar", "er", "ir"):
return None
vc, st, i = _vclass(lemma), _stem(lemma), _slot_idx(person, number)
if tense == "future":
return lemma + _FUT[i]
if tense == "conditional":
return lemma + _COND[i]
table = _REG.get((mood, tense, vc))
if table:
return st + table[i]
if mood == "imp" and tense == "present":
# affirmative tú imperative = 3sg present indicative
pres = _REG.get(("ind", "present", vc))
return st + pres[2] if number == "singular" else st + pres[5]
return None
# ── PUBLIC: verb conjugation ────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number, formality="informal"):
"""Return (surface, confidence). mood in ind|sbjv|imp; tense per _VERB_KEYMAP."""
lemma = lemma.strip().lower()
p, n = _PERSON.get(person), _NUMBER.get(number)
formal = "form" if formality == "formal" else "infm"
if p and n:
for fkey in (formal, "any", "infm" if formal == "form" else "form"):
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}|{fkey}"))
if form:
return form, "lexicon"
m = _mlc_conjugate(lemma, mood, tense, person, number)
if m:
return m, "model"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r:
return r, "rule"
return lemma, "fallback"
_IRREG_PART = { # guarantee the common irregular participles
"escribir": "escrito", "describir": "descrito", "abrir": "abierto",
"cubrir": "cubierto", "descubrir": "descubierto", "morir": "muerto",
"poner": "puesto", "ver": "visto", "volver": "vuelto", "devolver": "devuelto",
"hacer": "hecho", "deshacer": "deshecho", "decir": "dicho", "romper": "roto",
"resolver": "resuelto", "freír": "frito", "imprimir": "impreso",
"satisfacer": "satisfecho", "prever": "previsto", "revolver": "revuelto",
}
def participle(lemma):
lemma = lemma.strip().lower()
if lemma in _IRREG_PART:
return _IRREG_PART[lemma], "lexicon"
if lemma in _PART:
return _PART[lemma], "lexicon"
if lemma.endswith("ar"):
return lemma[:-2] + "ado", "rule"
if lemma[-2:] in ("er", "ir"):
return lemma[:-2] + "ido", "rule"
return lemma, "fallback"
_IRREG_GER = {"dormir": "durmiendo", "morir": "muriendo", "pedir": "pidiendo",
"sentir": "sintiendo", "mentir": "mintiendo", "servir": "sirviendo",
"venir": "viniendo", "decir": "diciendo", "poder": "pudiendo",
"ir": "yendo", "leer": "leyendo", "creer": "creyendo",
"oír": "oyendo", "traer": "trayendo", "caer": "cayendo",
"construir": "construyendo", "huir": "huyendo", "reír": "riendo"}
def gerund(lemma):
lemma = lemma.strip().lower()
if lemma in _IRREG_GER:
return _IRREG_GER[lemma], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
if lemma.endswith("ar"):
return lemma[:-2] + "ando", "rule"
if lemma[-2:] in ("er", "ir"):
return lemma[:-2] + "iendo", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ────────────────────────────────────────────────
_INVARIANT_PL = {"lunes", "martes", "miércoles", "jueves", "viernes",
"crisis", "tesis", "análisis", "dosis", "virus", "paraguas"}
def _gender_heuristic(noun):
for suf, g in (("ión", "f"), ("dad", "f"), ("tad", "f"), ("umbre", "f"),
("sis", "f"), ("ez", "f"), ("triz", "f"),
("ema", "m"), ("ama", "m"), ("oma", "m"), ("aje", "m"),
("or", "m"), ("án", "m"), ("ín", "m")):
if noun.endswith(suf):
return g
if noun.endswith("o"):
return "m"
if noun.endswith("a"):
return "f"
return "m"
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g"):
return d["g"]
return _gender_heuristic(lemma)
def _regular_plural(noun):
if noun in _INVARIANT_PL:
return noun
if not noun:
return noun
last = noun[-1]
if last == "z":
return noun[:-1] + "ces"
if last in "aeiouáéíóú":
# stressed final vowel í/ú -> +es (rubí->rubíes), else +s
if last in "íú":
return noun + "es"
return noun + "s"
if last == "s":
# esdrújula / stress-final handled crudely; most polysyllables invariant
return noun
return noun + "es"
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
num = "SG" if number == "singular" else "PL"
if d:
# honor a requested gender for animate nouns (gato -> gata)
if gender and (gender, num) in d:
return d[(gender, num)], "lexicon"
if d.get(num):
return d[num], "lexicon"
if number == "singular":
return lemma, "rule" if not d else "lexicon"
return _regular_plural(lemma), "rule"
# ── PUBLIC: adjective agreement ─────────────────────────────────────────────────
_INV_GENDER_ADJ = {"español": "española", "trabajador": "trabajadora",
"hablador": "habladora", "encantador": "encantadora",
"alemán": "alemana", "francés": "francesa", "inglés": "inglesa"}
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
d = _ADJS.get(lemma)
num = "SG" if number == "singular" else "PL"
if d:
form = d.get((gender, num))
if form:
return form, "lexicon"
# gender-invariant adjective (grande, feliz, azul): fem == masc.
# For a missing plural, pluralize this gender's singular form.
sg = d.get((gender, "SG")) or d.get(("m", "SG")) or lemma
if number == "plural":
return _regular_plural(sg), "rule"
return sg, "lexicon"
# rule fallback
a = lemma
if gender == "f":
if a in _INV_GENDER_ADJ:
a = _INV_GENDER_ADJ[a]
elif a.endswith("o"):
a = a[:-1] + "a"
if number == "plural":
a = _regular_plural(a)
return a, ("rule" if (a != lemma or gender == "m") else "rule")
# ── PUBLIC: clitic enclisis (dá + me + lo -> dámelo) ────────────────────────────
def _strip_accents(s):
return "".join(c for c in unicodedata.normalize("NFD", s)
if unicodedata.category(c) != "Mn")
def _count_syllables_vowelgroups(word):
# crude: count vowel groups
w = _strip_accents(word).lower()
groups, prev = 0, False
for ch in w:
isv = ch in "aeiou"
if isv and not prev:
groups += 1
prev = isv
return groups
def _host_stress_from_end(word):
"""Stressed-syllable index counted from the end (1=last) of a verb host."""
syls = _count_syllables_vowelgroups(word)
if any(c in "áéíóú" for c in word):
return None # already carries its own accent
if word[-2:] in ("ar", "er", "ir"): # infinitive: oxytone
return 1
if word.endswith("ndo"): # gerund: paroxytone
return 2
if word[-1:] in "aeiouns" and syls >= 2: # default paroxytone
return 2
return 1 # monosyllable / consonant-final oxytone
def attach_enclitics(verb_form, clitics):
"""Append clitic pronouns to a verb (imperative/infinitive/gerund enclisis)
and add a written accent when the resulting word becomes esdrújula/
sobreesdrújula (stress >= 3 syllables from the end): dá+me+lo -> dámelo,
lleva+me -> llévame, but dar+te -> darte and da+me -> dame (no accent)."""
if not clitics:
return verb_form
tail = "".join(clitics)
if any(c in "áéíóú" for c in verb_form): # host already accented
return verb_form + tail
sfe = _host_stress_from_end(verb_form)
total_sfe = sfe + len(clitics) # each clitic = 1 syllable
if total_sfe >= 3:
return _accentuate_nucleus(verb_form, sfe) + tail
return verb_form + tail
def _accentuate_nucleus(word, sfe):
"""Put a written accent on the syllable `sfe` positions from the word's end."""
vowels = "aeiou"
nuclei = [i for i, ch in enumerate(word) if ch in vowels]
if not nuclei or sfe > len(nuclei):
return word
i = nuclei[-sfe]
acc = {"a": "á", "e": "é", "i": "í", "o": "ó", "u": "ú"}
return word[:i] + acc[word[i]] + word[i + 1:]
def _accentuate_last_stressed(word):
# Restore the host's ORIGINAL lexical stress with a written accent.
# Default Spanish stress: word ending in vowel/n/s -> penultimate syllable;
# otherwise (e.g. infinitives in -r) -> last syllable.
vowels = "aeiou"
nuclei = [i for i, ch in enumerate(word) if ch in vowels]
if not nuclei:
return word
if word[-1] in "aeiouns" and len(nuclei) >= 2:
i = nuclei[-2] # paroxytone: penult nucleus
else:
i = nuclei[-1] # oxytone / monosyllable: last nucleus
acc = {"a": "á", "e": "é", "i": "í", "o": "ó", "u": "ú"}
return word[:i] + acc[word[i]] + word[i + 1:]
def lexicon_stats():
return {
"source": "UniMorph Spanish (github.com/unimorph/spa)",
"license": "CC-BY-SA 3.0 (Wiktionary-derived)",
"total_forms": sum(len(v) for v in (_VERBS, _NOUNS, _ADJS)) if False else None,
"verb_forms": len(_VERBS),
"verb_lemmas": len({k[0] for k in _VERBS}),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
"participles": len(_PART),
"gerunds": len(_GER),
}
if __name__ == "__main__":
import json
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("hablar", "ind", "present", "first", "singular", "hablo"),
("comer", "ind", "present", "third", "plural", "comen"),
("vivir", "ind", "present", "first", "plural", "vivimos"),
("ser", "ind", "present", "third", "singular", "es"),
("ir", "ind", "preterite", "first", "singular", "fui"),
("tener", "ind", "future", "first", "singular", "tendré"),
("hacer", "sbjv", "present", "first", "singular", "haga"),
("dormir", "ind", "present", "first", "singular", "duermo"),
("pensar", "sbjv", "present", "third", "singular", "piense"),
("dar", "ind", "preterite", "third", "singular", "dio"),
("poner", "ind", "conditional", "first", "singular", "pondría"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
if got == exp:
ok += 1
print(f" {flag}{lemma:8} {mood}/{tense} {per[:3]}.{num[:2]:3} -> {got:14} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" gender casa:", noun_gender("casa"), "| problema:", noun_gender("problema"),
"| agua:", noun_gender("agua"), "| mano:", noun_gender("mano"))
print(" plural: luz->", inflect_noun("luz", "plural"), "| rey->", inflect_noun("rey", "plural"))
print(" adj: rojo/f/pl->", inflect_adj("rojo", "f", "plural"),
"| feliz/m/pl->", inflect_adj("feliz", "m", "plural"),
"| grande/f/pl->", inflect_adj("grande", "f", "plural"))
print(" enclisis: da+[me,lo]->", attach_enclitics("da", ["me", "lo"]),
"| di+[me]->", attach_enclitics("di", ["me"]),
"| dar+[se,lo]->", attach_enclitics("dar", ["se", "lo"]))
-629
View File
@@ -1,629 +0,0 @@
"""morphology_fr_full.py — production-grade French morphological generator.
Same architecture as morphology_it_full.py (shared Romance engine); French-specific
data and rules swapped in. Backed by three real, Wiktionary-lineage sources:
VERBS
UniMorph French (github.com/unimorph/fra, CC-BY-SA 3.0)
7,535 verb lemmas × full paradigm, CLEAN orthography:
indicatif présent / imparfait (PST;IPFV) / passé simple (PST;PFV) /
futur, conditionnel (COND), subjonctif présent (SBJV;PRS) /
subjonctif imparfait (SBJV;PST), impératif (POS;IMP), infinitif (NFIN),
participe présent (V.CVB/V.PTCP;PRS), participe passé (V.PTCP;PST, m.sg).
fr_irreg_verbs.json — high-frequency verbs UniMorph MISSES or mis-slots,
above all ÊTRE (absent from UniMorph fra), plus avoir/aller/faire/… — the
auxiliaries the passé-composé + être-agreement system depends on. Extracted
from kaikki.org French (build_fr_irreg.py), reflexive/multiword forms
dropped. This layer takes PRIORITY.
NOUNS + ADJECTIVES — kaikki.org French (Wiktionary extract, CC-BY-SA 3.0)
noun lemmas WITH inherent gender (head-template arg) + real plural
(cheval->chevaux, œil->yeux, invariable -s/-x/-z), resolved PER LEMMA.
adjective lemmas with real feminine + plural (petit->petite/petits/petites,
beau->belle/beaux/belles, heureux->heureuse, rouge invariant-gender).
Fallbacks (degrade, never crash, on OOV input):
verbs : rule generator for -er / -ir(-iss-) / -re (with -cer/-ger spelling,
future/conditional stems, imparfait/subjonctif endings)
nouns : gender heuristic (endings) + rule pluralization (-al->-aux, -eau->-eaux)
adjs : fem/plural agreement rules (-er->-ère, -eux->-euse, -f->-ve, +e default)
Confidence flag on every form: "lexicon" | "rule" | "fallback".
Public API (used by realizer_fr.py): identical signature to morphology_it_full.
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "fra.unimorph")
_IRREG = os.path.join(_HERE, "data", "fr_irreg_verbs.json")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_fr.jsonl")
_CACHE = os.path.join(_HERE, "data", "fr_morph_cache.pkl")
# ── (mood, tense) -> UniMorph feature set that must ALL be present ────────────────
_VERB_KEYMAP = {
("ind", "present"): {"IND", "PRS"},
("ind", "imperfect"): {"IND", "PST", "IPFV"}, # imparfait
("ind", "passe_simple"): {"IND", "PST", "PFV"}, # passé simple
("ind", "future"): {"IND", "FUT"},
("ind", "conditional"): {"COND"}, # French: V;COND;1;SG
("sbjv", "present"): {"SBJV", "PRS"},
("sbjv", "imperfect"): {"SBJV", "PST"},
("imp", "affirmative"): {"POS", "IMP"},
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── build verb lexicon from UniMorph ─────────────────────────────────────────────
def _build_verbs():
verbs = {}
part = {}
ger = {}
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f:
part.setdefault(lemma, form)
elif "PRS" in f:
ger.setdefault(lemma, form)
continue
if head == "V.CVB":
if "PRS" in f:
ger.setdefault(lemma, form)
continue
if head != "V":
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
for (mood, tense), req in _VERB_KEYMAP.items():
if not req <= f:
continue
if tense == "imperfect" and "PFV" in f:
continue
if tense == "passe_simple" and "IPFV" in f:
continue
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), form)
break
return verbs, part, ger
# ── kaikki nouns + adjectives ────────────────────────────────────────────────────
_EXCL_FORM_TAGS = {"alternative", "archaic", "obsolete", "dialectal", "regional",
"diminutive", "augmentative", "pejorative", "comparative",
"superlative", "misspelling", "rare", "informal", "literary",
"poetic", "error-unrecognized-form", "construed", "collective",
"nonstandard", "dated", "Louisiana", "Switzerland", "Belgium"}
def _kaikki_gender(arg):
if not arg:
return None
a = str(arg).lower()
if a.startswith("f"):
return "f"
if a.startswith("m"):
return "m"
return None
def _build_nouns_adjs():
nouns = {}
adjs = {}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
pos = d.get("pos")
word = d.get("word", "")
if not word or " " in word:
continue
forms = d.get("forms", []) or []
if pos == "noun":
ht = d.get("head_templates") or []
g = None
if ht:
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
if g is None:
tags = d.get("tags") or []
if "feminine" in tags:
g = "f"
elif "masculine" in tags:
g = "m"
pl = None
for x in forms:
t = set(x.get("tags") or [])
if "plural" in t and not (t & _EXCL_FORM_TAGS):
fm = x.get("form")
if fm and " " not in fm and fm not in ("#", "-", ""):
pl = fm
break
if word not in nouns:
nouns[word] = {"g": g, "SG": word, "PL": pl}
else:
cur = nouns[word]
if cur.get("g") is None and g:
cur["g"] = g
if not cur.get("PL") and pl:
cur["PL"] = pl
elif pos == "adj":
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word)
for x in forms:
t = set(x.get("tags") or [])
fm = x.get("form")
if not fm or " " in fm or (t & _EXCL_FORM_TAGS):
continue
if "feminine" in t and "plural" in t:
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
elif "masculine" in t and "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
elif "feminine" in t:
d0[("f", "SG")] = d0.get(("f", "SG")) or fm
elif "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
return nouns, adjs
def _build_cache():
verbs, part, ger = _build_verbs()
nouns, adjs = _build_nouns_adjs()
with open(_IRREG, encoding="utf-8") as fh:
irreg = json.load(fh)
data = {"verbs": verbs, "part": part, "ger": ger,
"nouns": nouns, "adjs": adjs, "irreg": irreg}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [_UNIMORPH, _KAIKKI, _IRREG]
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PART, _GER, _NOUNS, _ADJS, _IRREGV = (
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"],
_LEX["irreg"])
# ── regular-ending rule fallback ─────────────────────────────────────────────────
def _vclass(lemma):
if lemma.endswith("er"):
return "er"
if lemma.endswith("ir"):
return "ir"
if lemma.endswith("re"):
return "re"
if lemma.endswith("oir"):
return "oir"
return None
# present-tense endings [1sg,2sg,3sg,1pl,2pl,3pl]
_REG_PRES = {
"er": ["e", "es", "e", "ons", "ez", "ent"],
"ir": ["is", "is", "it", "issons", "issez", "issent"], # -iss- class (finir)
"re": ["s", "s", "", "ons", "ez", "ent"], # vendre: vends/vend
}
_REG_IMPF = ["ais", "ais", "ait", "ions", "iez", "aient"] # attaches to pres-1pl stem
_REG_SUBJ = ["e", "es", "e", "ions", "iez", "ent"] # attaches to 3pl stem
_REG_PS = { # passé simple
"er": ["ai", "as", "a", "âmes", "âtes", "èrent"],
"ir": ["is", "is", "it", "îmes", "îtes", "irent"],
"re": ["is", "is", "it", "îmes", "îtes", "irent"],
}
_FUT = ["ai", "as", "a", "ons", "ez", "ont"]
_COND = ["ais", "ais", "ait", "ions", "iez", "aient"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _fut_stem(lemma, vc):
"""Future/conditional stem = infinitive (drop final -e of -re)."""
if vc == "re":
return lemma[:-1] # vendre -> vendr-
return lemma # parler-, finir-
def _pres_1pl_stem(lemma, vc):
"""Imparfait stem = present 1pl minus -ons (parlons->parl-, finissons->finiss-)."""
if vc == "er":
stem = lemma[:-2]
if stem.endswith("g"):
return stem + "e" # mangeons -> mange- (imparfait mangeais)
if stem.endswith("c"):
return stem[:-1] + "ç" # commençons -> commenç-
return stem
if vc == "ir":
return lemma[:-1] + "iss" # finir -> finiss-
if vc == "re":
return lemma[:-2] # vendre -> vend-
return lemma[:-2]
def _apply_er_spelling(stem, ending):
"""-cer/-ger softening before a/o (commençons, mangeons)."""
if ending and ending[0] in ("a", "o"):
if stem.endswith("c"):
return stem[:-1] + "ç" + ending
if stem.endswith("g"):
return stem + "e" + ending
return stem + ending
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
i = _slot_idx(person, number)
if mood == "ind" and tense in ("future", "conditional"):
stem = _fut_stem(lemma, vc)
end = (_FUT if tense == "future" else _COND)[i]
return stem + end
if mood == "ind" and tense == "present":
table = _REG_PRES.get("ir" if vc == "ir" else vc)
if not table:
return None
body = lemma[:-2] if vc in ("er", "re") else lemma[:-1] if vc == "ir" else lemma[:-2]
if vc == "ir":
body = lemma[:-2] # fin- ; endings carry -iss-
end = table[i]
return body + end
end = table[i]
if vc == "er":
return _apply_er_spelling(body, end)
return body + end
if mood == "ind" and tense == "imperfect":
stem = _pres_1pl_stem(lemma, vc)
return stem + _REG_IMPF[i]
if mood == "ind" and tense == "passe_simple":
table = _REG_PS.get("ir" if vc == "ir" else vc)
if not table:
return None
body = lemma[:-2] if vc in ("er", "re") else lemma[:-2]
end = table[i]
if vc == "er":
return _apply_er_spelling(body, end)
return body + end
if mood == "sbjv" and tense == "present":
# subjonctif: present-3pl stem + e/es/e/ions/iez/ent
stem3 = _pres_1pl_stem(lemma, vc) if vc == "ir" else (
lemma[:-2] if vc in ("er", "re") else lemma[:-2])
if vc == "ir":
stem3 = lemma[:-2] + "iss"
end = _REG_SUBJ[i]
if vc == "er":
return _apply_er_spelling(stem3, end)
return stem3 + end
if mood == "imp" and tense == "affirmative":
# impératif ~ present indicative (tu drops -s for -er verbs)
pres = _rule_conjugate(lemma, "ind", "present", person, number)
if pres and vc == "er" and person == "second" and number == "singular":
return pres[:-1] if pres.endswith("es") else pres
return pres
return None
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
"""Return (surface, confidence)."""
lemma = lemma.strip().lower()
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{number}"
ir = _IRREGV.get(lemma)
if ir and key in ir:
return ir[key], "lexicon"
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
if form:
return form, "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r:
return r, "rule"
return lemma, "fallback"
# ── PUBLIC: participle + gerund/participe présent ────────────────────────────────
def _participle_msg(lemma):
ir = _IRREGV.get(lemma)
if ir and "part" in ir:
return ir["part"], "lexicon"
if lemma in _PART:
return _PART[lemma], "lexicon"
return None, None
# irregular participle fem/plural quirks (drop circonflexe: dû->due, dus)
_PART_FIX = {"": {"f|SG": "due", "m|PL": "dus", "f|PL": "dues"}}
def participle(lemma, gender="m", number="singular"):
"""Past participle with French gender/number agreement.
m.sg = base; f.sg = base+e; m.pl = base+s (invariable if base ends s/x);
f.pl = f.sg+s."""
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
msg, src = _participle_msg(lemma)
conf = "lexicon"
if msg is None:
vc = _vclass(lemma)
if vc == "er":
msg = lemma[:-2] + "é"
elif vc == "ir":
msg = lemma[:-1] # finir -> fini, partir -> parti
elif vc == "re":
msg = lemma[:-2] + "u" # vendre -> vendu
elif vc == "oir":
msg = lemma[:-3] + "u" # (rough) recevoir handled by irreg
else:
return lemma, "fallback"
conf = "rule"
fix = _PART_FIX.get(msg)
if fix and f"{g}|{num}" in fix:
return fix[f"{g}|{num}"], conf
if g == "m" and num == "SG":
return msg, conf
fem = msg + "e" if not msg.endswith("e") else msg
if g == "f" and num == "SG":
return fem, conf
if g == "m" and num == "PL":
return msg if msg.endswith(("s", "x")) else msg + "s", conf
# f|PL
return fem + "s", conf
def gerund(lemma):
"""Participe présent (base for gérondif 'en -ant')."""
lemma = lemma.strip().lower()
ir = _IRREGV.get(lemma)
if ir and "ger" in ir:
return ir["ger"], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
vc = _vclass(lemma)
if vc == "er":
stem = lemma[:-2]
if stem.endswith("g"):
return stem + "eant", "rule"
if stem.endswith("c"):
return stem[:-1] + "çant", "rule"
return stem + "ant", "rule"
if vc == "ir":
return lemma[:-2] + "issant", "rule"
if vc == "re":
return lemma[:-2] + "ant", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
_FEM_SUF = ("tion", "sion", "aison", "ance", "ence", "ette", "elle", "esse",
"ude", "ade", "ée", "", "tié", "ie", "ise", "ure", "eur")
_MASC_SUF = ("ment", "age", "eau", "isme", "oir", "ier", "eur", "in", "on")
def _gender_heuristic(noun):
for suf in _FEM_SUF:
if noun.endswith(suf):
return "f"
for suf in _MASC_SUF:
if noun.endswith(suf):
return "m"
if noun.endswith("e"):
return "f"
return "m"
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g") in ("m", "f"):
return d["g"]
return _gender_heuristic(lemma)
# closed sets for French plural irregularities
_OU_X = {"bijou", "caillou", "chou", "genou", "hibou", "joujou", "pou"}
_AIL_AUX = {"travail", "vitrail", "corail", "émail", "bail", "soupirail", "vantail"}
_AL_S = {"bal", "carnaval", "festival", "récital", "chacal", "régal", "cal", "aval"}
def _rule_plural(noun, gender):
"""Deterministic French pluralization. (form, ok); ok=False FLAGS ambiguity."""
if not noun:
return noun, True
if noun[-1:] in ("s", "x", "z"):
return noun, True # invariable
if noun in _OU_X:
return noun + "x", True
if noun.endswith(("eau", "au", "eu")):
if noun in ("pneu", "bleu", "landau", "sarrau"):
return noun + "s", True
return noun + "x", True # bateau->bateaux, jeu->jeux
if noun.endswith("al"):
if noun in _AL_S:
return noun + "s", True
return noun[:-2] + "aux", True # cheval->chevaux
if noun.endswith("ail"):
if noun in _AIL_AUX:
return noun[:-3] + "aux", True # travail->travaux
return noun + "s", True
return noun + "s", True # default
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if number == "singular":
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
if d and d.get("PL"):
return d["PL"], "lexicon"
g = gender or noun_gender(lemma)
form, ok = _rule_plural(lemma, g)
return form, ("rule" if ok else "fallback")
# adjectives whose kaikki entries are unreliable: audited forms
_ADJ_FIX = {
"beau": {("m", "SG"): "beau", ("f", "SG"): "belle",
("m", "PL"): "beaux", ("f", "PL"): "belles"},
"nouveau": {("m", "SG"): "nouveau", ("f", "SG"): "nouvelle",
("m", "PL"): "nouveaux", ("f", "PL"): "nouvelles"},
"vieux": {("m", "SG"): "vieux", ("f", "SG"): "vieille",
("m", "PL"): "vieux", ("f", "PL"): "vieilles"},
"fou": {("m", "SG"): "fou", ("f", "SG"): "folle",
("m", "PL"): "fous", ("f", "PL"): "folles"},
"blanc": {("m", "SG"): "blanc", ("f", "SG"): "blanche",
("m", "PL"): "blancs", ("f", "PL"): "blanches"},
"long": {("m", "SG"): "long", ("f", "SG"): "longue",
("m", "PL"): "longs", ("f", "PL"): "longues"},
"bon": {("m", "SG"): "bon", ("f", "SG"): "bonne",
("m", "PL"): "bons", ("f", "PL"): "bonnes"},
}
def _rule_fem(a):
if a.endswith("e"):
return a
if a.endswith("er"):
return a[:-2] + "ère"
if a.endswith("eau"):
return a[:-3] + "elle"
if a.endswith("eux"):
return a[:-3] + "euse"
if a.endswith("f"):
return a[:-1] + "ve"
if a.endswith(("on", "en", "el", "eil", "et")):
return a + a[-1] + "e" # bon->bonne, ancien->ancienne, muet->muette
if a.endswith("c"):
return a[:-1] + "che" # blanc->blanche (public->publique via FIX)
return a + "e" # grand->grande, petit->petite, vert->verte
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
fix = _ADJ_FIX.get(lemma)
if fix and (g, num) in fix:
return fix[(g, num)], "lexicon"
d = _ADJS.get(lemma)
if d and d.get((g, num)):
return d[(g, num)], "lexicon"
# derive
msc = (d.get(("m", "SG")) if d else None) or lemma
if g == "m" and num == "SG":
return msc, "lexicon" if d else "rule"
fem = (d.get(("f", "SG")) if d else None) or _rule_fem(msc)
if g == "f" and num == "SG":
return fem, "lexicon" if (d and d.get(("f", "SG"))) else "rule"
if g == "m" and num == "PL":
if msc.endswith(("s", "x")):
return msc, "rule"
if msc.endswith("al"):
return msc[:-2] + "aux", "rule"
if msc.endswith("eau"):
return msc + "x", "rule"
return msc + "s", "rule"
# f|PL
return (fem if fem.endswith("s") else fem + "s"), "rule"
def lexicon_stats():
return {
"verb_source": "UniMorph French (github.com/unimorph/fra) + kaikki.org "
"irregulars (être + high-frequency)",
"noun_adj_source": "kaikki.org French (Wiktionary extract)",
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
"unimorph_verb_forms": len(_VERBS),
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
"irregular_verb_lemmas": len(_IRREGV),
"participle_lemmas": len(_PART),
"gerund_lemmas": len(_GER),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("parler", "ind", "present", "first", "singular", "parle"),
("être", "ind", "present", "third", "singular", "est"),
("avoir", "ind", "present", "first", "singular", "ai"),
("aller", "ind", "present", "third", "plural", "vont"),
("finir", "ind", "present", "first", "singular", "finis"),
("finir", "ind", "present", "first", "plural", "finissons"),
("manger", "ind", "present", "first", "plural", "mangeons"),
("faire", "ind", "future", "first", "singular", "ferai"),
("pouvoir", "sbjv", "present", "third", "singular", "puisse"),
("prendre", "ind", "passe_simple", "third", "singular", "prit"),
("vendre", "ind", "present", "third", "singular", "vend"),
("commencer", "ind", "imperfect", "first", "singular", "commençais"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
ok += got == exp
print(f" {flag}{lemma:10} {mood}/{tense:12} {per[:3]}.{num[:2]} -> {got:12} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" gender: maison=", noun_gender("maison"), "chat=", noun_gender("chat"),
"cheval=", noun_gender("cheval"), "nation=", noun_gender("nation"))
print(" plural: cheval->", inflect_noun("cheval", "plural"),
"| bateau->", inflect_noun("bateau", "plural"),
"| prix->", inflect_noun("prix", "plural"),
"| chat->", inflect_noun("chat", "plural"))
print(" adj: petit/f/sg->", inflect_adj("petit", "f", "singular"),
"| beau/f/sg->", inflect_adj("beau", "f", "singular"),
"| heureux/f/sg->", inflect_adj("heureux", "f", "singular"),
"| national/m/pl->", inflect_adj("national", "m", "plural"))
print(" part: aller/f/sg->", participle("aller", "f", "singular"),
"| prendre/f/pl->", participle("prendre", "f", "plural"),
"| finir/m/pl->", participle("finir", "m", "plural"))
print(" ger: manger->", gerund("manger"), "| finir->", gerund("finir"))
-588
View File
@@ -1,588 +0,0 @@
"""morphology_it_full.py — production-grade Italian morphological generator.
NOT a toy. Backed by three real, Wiktionary-lineage lexical sources:
VERBS
UniMorph Italian (github.com/unimorph/ita, CC-BY-SA 3.0)
10,009 verb lemmas × full paradigm, CLEAN orthography (no stress marks):
indicative present / imperfetto (PST;IPFV) / passato remoto (PST;PFV) /
futuro, condizionale (COND),
congiuntivo presente (SBJV;PRS) / imperfetto (SBJV;PST),
affirmative imperative, infinitive, gerundio (V.CVB;PRS),
past participle (masc-sg; fem/plural derived by vowel rule).
it_irreg_verbs.json — 66 high-frequency verbs UniMorph MISSES
(essere, avere, potere, uscire, tenere, prendere, piacere, …), extracted
from kaikki.org Italian, filtered to standard forms, and DE-STRESSED to
real orthography (kaikki marks tonic stress everywhere: pàrlo->parlo,
avùto->avuto; final legit accents kept: sarò, è). Built by build_it_irreg.py.
This layer takes priority — it supplies the two auxiliaries essere/avere,
which the whole passato-prossimo / essere-agreement system depends on.
NOUNS + ADJECTIVES — kaikki.org Italian (Wiktionary extract, CC-BY-SA 3.0)
noun lemmas WITH inherent gender (head-template arg) + real (often irregular)
plural — uomo->uomini, uovo->uova, dito->dita, città invariant — resolved
PER LEMMA, never guessed.
adjective lemmas with real feminine + masc/fem plural (italiano->italiana/
italiani/italiane, felice->felici invariant).
Fallbacks (degrade, never crash, on OOV input):
verbs : rule generator for regular -are/-ere/-ire (with -care/-gare h-insertion
and -ciare/-giare/-iare i-drop spelling rules)
nouns : gender heuristic (endings) + rule pluralization (ambiguous -co/-go FLAGGED)
adjs : -o/-a/-e gender rule + rule pluralization
Confidence flag on every form:
"lexicon" from UniMorph / kaikki-irregular / kaikki noun-adj (trust: high)
"rule" deterministic rule (trust: medium)
"fallback" could not inflect; returned lemma / ambiguous (trust: low -> FLAG)
Public API (used by realizer_it.py):
conjugate(lemma, mood, tense, person, number) -> (form, conf)
participle(lemma, gender="m", number="singular") -> (form, conf)
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"
inflect_noun(lemma, number, gender=None) -> (form, conf)
inflect_adj(lemma, gender, number) -> (form, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "ita.unimorph")
_IRREG = os.path.join(_HERE, "data", "it_irreg_verbs.json")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_it.jsonl")
_CACHE = os.path.join(_HERE, "data", "it_morph_cache.pkl")
# ── (mood, tense) -> UniMorph feature set that must ALL be present ────────────────
_VERB_KEYMAP = {
("ind", "present"): {"IND", "PRS"},
("ind", "imperfect"): {"IND", "PST", "IPFV"},
("ind", "passato_remoto"): {"IND", "PST", "PFV"},
("ind", "future"): {"IND", "FUT"},
("ind", "conditional"): {"COND"},
("sbjv", "present"): {"SBJV", "PRS"},
("sbjv", "imperfect"): {"SBJV", "PST"},
("imp", "affirmative"): {"POS", "IMP"},
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── build verb lexicon from UniMorph ─────────────────────────────────────────────
def _build_verbs():
verbs = {} # (lemma, "mood|tense|person|number") -> form
part = {} # lemma -> masc-sg past participle
ger = {} # lemma -> gerundio
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f:
part.setdefault(lemma, form)
continue
if head == "V.CVB": # gerundio (converb, present)
if "PRS" in f:
ger.setdefault(lemma, form)
continue
if head != "V":
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
for (mood, tense), req in _VERB_KEYMAP.items():
# exact-set discipline: PST;PFV must not match PST;IPFV, etc.
if not req <= f:
continue
# guard IND;PST ambiguity: require the specific aspect feature
if tense == "imperfect" and "PFV" in f:
continue
if tense == "passato_remoto" and "IPFV" in f:
continue
# COND must not also be a subjunctive/imperative slot
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), form)
break
return verbs, part, ger
# ── kaikki nouns + adjectives ────────────────────────────────────────────────────
_EXCL_FORM_TAGS = {"alternative", "archaic", "obsolete", "dialectal", "regional",
"diminutive", "augmentative", "pejorative", "comparative",
"superlative", "misspelling", "rare", "informal", "literary",
"poetic", "error-unrecognized-form", "apocopic", "obsolete",
"construed", "collective"}
def _kaikki_gender(arg):
if not arg:
return None
a = str(arg).lower()
if a.startswith("f"):
return "f"
if a.startswith("m"):
return "m"
return None
def _build_nouns_adjs():
nouns = {} # lemma -> {"g","SG","PL"}
adjs = {} # lemma -> {("m","SG"),("f","SG"),("m","PL"),("f","PL")}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
pos = d.get("pos")
word = d.get("word", "")
if not word or " " in word:
continue
forms = d.get("forms", []) or []
if pos == "noun":
ht = d.get("head_templates") or []
g = None
if ht:
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
if g is None:
tags = d.get("tags") or []
if "feminine" in tags:
g = "f"
elif "masculine" in tags:
g = "m"
pl = None
for x in forms:
t = set(x.get("tags") or [])
if "plural" in t and not (t & _EXCL_FORM_TAGS):
fm = x.get("form")
if fm and " " not in fm and fm != "#":
pl = fm
break
if word not in nouns:
nouns[word] = {"g": g, "SG": word, "PL": pl}
else:
cur = nouns[word]
if cur.get("g") is None and g:
cur["g"] = g
if not cur.get("PL") and pl:
cur["PL"] = pl
elif pos == "adj":
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word)
for x in forms:
t = set(x.get("tags") or [])
fm = x.get("form")
if not fm or " " in fm or (t & _EXCL_FORM_TAGS):
continue
if "feminine" in t and "plural" in t:
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
elif "masculine" in t and "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
elif "feminine" in t:
d0[("f", "SG")] = d0.get(("f", "SG")) or fm
elif "plural" in t: # invariant-gender adj (felice -> felici)
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
return nouns, adjs
def _build_cache():
verbs, part, ger = _build_verbs()
nouns, adjs = _build_nouns_adjs()
with open(_IRREG, encoding="utf-8") as fh:
irreg = json.load(fh)
data = {"verbs": verbs, "part": part, "ger": ger,
"nouns": nouns, "adjs": adjs, "irreg": irreg}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [_UNIMORPH, _KAIKKI, _IRREG]
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PART, _GER, _NOUNS, _ADJS, _IRREGV = (
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"],
_LEX["irreg"])
# ── regular-ending rule fallback ─────────────────────────────────────────────────
def _vclass(lemma):
if lemma.endswith("are"):
return "are"
if lemma.endswith("ere"):
return "ere"
if lemma.endswith("ire"):
return "ire"
return None
# endings [1sg,2sg,3sg,1pl,2pl,3pl]
_REG = {
("ind", "present", "are"): ["o", "i", "a", "iamo", "ate", "ano"],
("ind", "present", "ere"): ["o", "i", "e", "iamo", "ete", "ono"],
("ind", "present", "ire"): ["o", "i", "e", "iamo", "ite", "ono"],
("ind", "imperfect", "are"): ["avo", "avi", "ava", "avamo", "avate", "avano"],
("ind", "imperfect", "ere"): ["evo", "evi", "eva", "evamo", "evate", "evano"],
("ind", "imperfect", "ire"): ["ivo", "ivi", "iva", "ivamo", "ivate", "ivano"],
("ind", "passato_remoto", "are"): ["ai", "asti", "ò", "ammo", "aste", "arono"],
("ind", "passato_remoto", "ere"): ["ei", "esti", "é", "emmo", "este", "erono"],
("ind", "passato_remoto", "ire"): ["ii", "isti", "ì", "immo", "iste", "irono"],
("sbjv", "present", "are"): ["i", "i", "i", "iamo", "iate", "ino"],
("sbjv", "present", "ere"): ["a", "a", "a", "iamo", "iate", "ano"],
("sbjv", "present", "ire"): ["a", "a", "a", "iamo", "iate", "ano"],
("sbjv", "imperfect", "are"): ["assi", "assi", "asse", "assimo", "aste", "assero"],
("sbjv", "imperfect", "ere"): ["essi", "essi", "esse", "essimo", "este", "essero"],
("sbjv", "imperfect", "ire"): ["issi", "issi", "isse", "issimo", "iste", "issero"],
# imperative: 2sg,3sg(Lei),1pl,2pl,3pl (1sg has none)
("imp", "affirmative", "are"): [None, "a", "i", "iamo", "ate", "ino"],
("imp", "affirmative", "ere"): [None, "i", "a", "iamo", "ete", "ano"],
("imp", "affirmative", "ire"): [None, "i", "a", "iamo", "ite", "ano"],
}
# future / conditional attach to a stem = infinitive minus final -e, with
# -are -> -er (parlare->parler-), -ere/-ire keep (credere->creder-, dormir-)
_FUT = ["ò", "ai", "à", "emo", "ete", "anno"]
_COND = ["ei", "esti", "ebbe", "emmo", "este", "ebbero"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _fut_stem(lemma, vc):
body = lemma[:-3] # drop are/ere/ire
if vc == "are":
return body + "er"
return body + vc[0] + "r" # ere->er? no: keep vowel: creder-, dormir-
# NOTE corrected below
def _apply_are_spelling(stem, ending):
"""-care/-gare insert h before front endings; -ciare/-giare/-sciare/-iare drop i."""
front = ending[:1] in ("i", "e")
if stem.endswith(("c", "g")) and front:
return stem + "h" + ending
if stem.endswith(("ci", "gi", "sci")) and ending[:1] == "i":
return stem[:-1] + ending # mangi+iamo -> mangiamo
if stem.endswith("i") and ending[:1] == "i":
return stem[:-1] + ending # studi+iamo -> studiamo
return stem + ending
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
body = lemma[:-3]
i = _slot_idx(person, number)
if mood == "ind" and tense in ("future", "conditional"):
stem = body + "er" if vc == "are" else body + vc[0] + "r"
# ere: creder-, ire: dormir- -> body + 'e'/'i' + 'r'
if vc == "ere":
stem = body + "er"
elif vc == "ire":
stem = body + "ir"
end = (_FUT if tense == "future" else _COND)[i]
# spelling: -care/-gare -> cherò/gherò ; -ciare/-giare -> cerò/gerò
if vc == "are":
if body.endswith(("c", "g")):
stem = body + "her"
elif body.endswith(("ci", "gi", "sci")):
stem = body[:-1] + "er"
elif body.endswith("i"):
stem = body[:-1] + "er"
return stem + end
table = _REG.get((mood, tense, vc))
if not table:
return None
end = table[i]
if end is None:
return None
if vc == "are":
return _apply_are_spelling(body, end)
# -ere/-ire: guard against double-i (dormi+iamo -> dormiamo)
if body.endswith("i") and end[:1] == "i":
return body[:-1] + end
return body + end
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
"""Return (surface, confidence). mood in ind|sbjv|imp; tense per _VERB_KEYMAP."""
lemma = lemma.strip().lower()
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{number}"
ir = _IRREGV.get(lemma)
if ir and key in ir:
return ir[key], "lexicon"
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
if form:
return form, "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r:
return r, "rule"
return lemma, "fallback"
# ── PUBLIC: participle + gerund ──────────────────────────────────────────────────
def _participle_msg(lemma):
"""Return (masc-sg participle, source) or (None, None)."""
ir = _IRREGV.get(lemma)
if ir and "part" in ir:
return ir["part"], "lexicon"
if lemma in _PART:
return _PART[lemma], "lexicon"
return None, None
def participle(lemma, gender="m", number="singular"):
"""Past participle with gender/number agreement (for essere-perfect & passives).
UniMorph/irregular give masc-sg; fem/plural derived by final-vowel swap
(-o -> -a/-i/-e), valid for regular -ato/-uto/-ito AND irregulars
(preso->presa/presi/prese, aperto->aperta/aperti/aperte, morto->morta/...)."""
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
msg, src = _participle_msg(lemma)
conf = "lexicon"
if msg is None:
vc = _vclass(lemma)
if vc == "are":
msg = lemma[:-3] + "ato"
elif vc == "ere":
msg = lemma[:-3] + "uto"
elif vc == "ire":
msg = lemma[:-3] + "ito"
else:
return lemma, "fallback"
conf = "rule"
# agreement: only -o participles inflect for gender+number
if msg.endswith("o"):
stem = msg[:-1]
suf = {"m|SG": "o", "f|SG": "a", "m|PL": "i", "f|PL": "e"}[f"{g}|{num}"]
return stem + suf, conf
return msg, conf # non -o participle: leave as-is (rare)
def gerund(lemma):
lemma = lemma.strip().lower()
ir = _IRREGV.get(lemma)
if ir and "ger" in ir:
return ir["ger"], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
vc = _vclass(lemma)
if vc == "are":
return lemma[:-3] + "ando", "rule"
if vc in ("ere", "ire"):
return lemma[:-3] + "endo", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
_FEM_SUF = ("zione", "sione", "gione", "", "", "trice", "aggine", "udine",
"igine", "ie", "essa", "izia", "ezza")
_MASC_SUF = ("ore", "ame", "iere", "ale", "ile")
def _gender_heuristic(noun):
for suf in _FEM_SUF:
if noun.endswith(suf):
return "f"
for suf in _MASC_SUF:
if noun.endswith(suf):
return "m"
if noun.endswith("o"):
return "m"
if noun.endswith("a"):
return "f"
if noun.endswith("à") or noun.endswith("ù"):
return "f"
return "m" # -e and consonant-final loanwords default masculine
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g") in ("m", "f"):
return d["g"]
return _gender_heuristic(lemma)
def _rule_plural(noun, gender):
"""Deterministic Italian pluralization. Returns (form, ok); ok=False FLAGS an
ambiguous case the lexicon would normally resolve (-co/-go palatalization)."""
if not noun:
return noun, True
# invariant: accented final vowel, consonant-final, monosyllable, -i final
if noun[-1:] in ("à", "è", "é", "ì", "í", "ò", "ó", "ù", "ú"):
return noun, True
if noun[-1:] not in ("a", "e", "o", "i", "u"):
return noun, True # consonant-final loanword: invariant
if noun.endswith("i"):
return noun, True # e.g. crisi, analisi: invariant
if noun.endswith("io"):
return noun[:-2] + "i", True # figlio->figli (unstressed i)
if noun.endswith("cia") or noun.endswith("gia"):
# vowel before cia/gia -> -cie/-gie ; consonant -> -ce/-ge (approx)
return noun[:-2] + "e", True # arancia->arance (majority)
if noun.endswith("ca"):
return noun[:-2] + "che", True # amica->amiche
if noun.endswith("ga"):
return noun[:-2] + "ghe", True
if noun.endswith("co"):
return noun[:-2] + "chi", False # AMBIGUOUS (amico->amici) -> flag
if noun.endswith("go"):
return noun[:-2] + "ghi", False # AMBIGUOUS (psicologo->psicologi)
if noun.endswith("a"):
return noun[:-1] + "e", True # casa->case (m -a: -i, but rare)
if noun.endswith("o"):
return noun[:-1] + "i", True # libro->libri
if noun.endswith("e"):
return noun[:-1] + "i", True # cane->cani, chiave->chiavi
return noun, True
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if number == "singular":
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
if d and d.get("PL"):
return d["PL"], "lexicon"
g = gender or noun_gender(lemma)
form, ok = _rule_plural(lemma, g)
return form, ("rule" if ok else "fallback")
# adjectives whose kaikki entries are unreliable (messy inflection templates):
# supply audited regular agreement forms (prenominal apocope handled in realizer).
_ADJ_FIX = {
"bello": {("m", "SG"): "bello", ("f", "SG"): "bella",
("m", "PL"): "belli", ("f", "PL"): "belle"},
"quello": {("m", "SG"): "quello", ("f", "SG"): "quella",
("m", "PL"): "quelli", ("f", "PL"): "quelle"},
}
# ── PUBLIC: adjective agreement ──────────────────────────────────────────────────
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
fix = _ADJ_FIX.get(lemma)
if fix and (g, num) in fix:
return fix[(g, num)], "lexicon"
d = _ADJS.get(lemma)
if d:
form = d.get((g, num))
if form:
return form, "lexicon"
sg = d.get((g, "SG")) or d.get(("m", "SG")) or lemma
if num == "PL":
pl, ok = _rule_plural(sg, g)
return pl, ("rule" if ok else "fallback")
return sg, "lexicon"
# rule fallback
a = lemma
if a.endswith("o"): # -o/-a/-i/-e class
base = a[:-1]
suf = {"m|SG": "o", "f|SG": "a", "m|PL": "i", "f|PL": "e"}[f"{g}|{num}"]
return base + suf, "rule"
if a.endswith("e"): # felice-class: SG invariant, PL -i
if num == "PL":
return a[:-1] + "i", "rule"
return a, "rule"
if num == "PL":
p, ok = _rule_plural(a, g)
return p, ("rule" if ok else "fallback")
return a, "rule"
def lexicon_stats():
return {
"verb_source": "UniMorph Italian (github.com/unimorph/ita) + kaikki.org "
"irregulars (de-stressed)",
"noun_adj_source": "kaikki.org Italian (Wiktionary extract)",
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
"unimorph_verb_forms": len(_VERBS),
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
"irregular_verb_lemmas": len(_IRREGV),
"participle_lemmas": len(_PART),
"gerund_lemmas": len(_GER),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("parlare", "ind", "present", "first", "singular", "parlo"),
("essere", "ind", "present", "third", "singular", "è"),
("avere", "ind", "present", "first", "singular", "ho"),
("mangiare", "ind", "present", "second", "singular", "mangi"),
("finire", "ind", "present", "first", "singular", "finisco"),
("andare", "ind", "present", "third", "plural", "vanno"),
("fare", "ind", "future", "first", "singular", "farò"),
("potere", "sbjv", "present", "third", "singular", "possa"),
("prendere", "ind", "passato_remoto", "first", "singular", "presi"),
("cercare", "ind", "present", "second", "singular", "cerchi"),
("dormire", "ind", "present", "third", "plural", "dormono"),
("credere", "ind", "future", "first", "singular", "crederò"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
ok += got == exp
print(f" {flag}{lemma:9} {mood}/{tense:14} {per[:3]}.{num[:2]} -> {got:12} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" gender: casa=", noun_gender("casa"), "problema=", noun_gender("problema"),
"mano=", noun_gender("mano"), "città=", noun_gender("città"),
"cane=", noun_gender("cane"))
print(" plural: uomo->", inflect_noun("uomo", "plural"),
"| uovo->", inflect_noun("uovo", "plural"),
"| città->", inflect_noun("città", "plural"),
"| amico->", inflect_noun("amico", "plural"),
"| casa->", inflect_noun("casa", "plural"))
print(" adj: italiano/f/pl->", inflect_adj("italiano", "f", "plural"),
"| felice/m/pl->", inflect_adj("felice", "m", "plural"),
"| bello/f/sg->", inflect_adj("bello", "f", "singular"))
print(" part: aprire/f/sg->", participle("aprire", "f", "singular"),
"| prendere/m/pl->", participle("prendere", "m", "plural"),
"| andare/f/sg->", participle("andare", "f", "singular"))
print(" ger: fare->", gerund("fare"), "| parlare->", gerund("parlare"))
-666
View File
@@ -1,666 +0,0 @@
# -*- coding: utf-8 -*-
"""morphology_lat_full.py — production-grade Latin morphological generator.
Latin is the FLAGSHIP dead-language realizer. It rides the *architecture* of the
Romance/Italic engine (the same Realization / spec-driven design and the UniMorph
loader pattern from morphology_it_full.py) but with the CASE SYSTEM RESTORED
the feature Romance lost. Latin therefore exercises machinery the modern Romance
siblings never needed: 5 declensions x 6 cases x 2 numbers x 3 genders, plus a
4-conjugation verb system with tense/mood/voice.
DATA (real, attested no fabrication):
NOUNS + ADJECTIVES UniMorph Latin (github.com/unimorph/lat, CC-BY-SA 3.0)
163,182 N forms across ~thousands of lemmas, each with the full case paradigm
N;NOM/GEN/DAT/ACC/ABL/VOC;SG/PL (real inflected forms, WITH macrons:
puella->puellam, rēx->rēgis, corpus->corporis).
244,197 ADJ forms with case x GENDER x number, incl. UniMorph's combined
tags (GEN+DAT, MASC+FEM, MASC+FEM+NEUT) which are split on load.
462,668 V.PTCP forms (participles) also carry case/gender/number.
UniMorph N tags DO NOT encode inherent gender, so noun gender is inferred
from the declension (nom-sg + gen-sg endings) with a curated exceptions
map the standard, attestable rule (1st decl -a/-ae = fem, 2nd -us/-i =
masc, -um = neut, ...).
VERBS RULE ENGINE (honest gap: UniMorph Latin's verb list is a 947-lemma
sample of rare/prefixed verbs that MISSES every core textbook verb amō,
videō, sum, regō, ... are all absent). Latin conjugation is, however, highly
regular, so verbs are generated by a deterministic 4-conjugation engine over
curated principal parts (present / perfect / supine stems), sourced from
standard references. Irregulars (sum, possum, , ferō, volō, nōlō, mālō)
are curated full tables. Forms are flagged "rule" (not "lexicon") for honesty.
Confidence flag on every form (same contract as the Romance engine):
"lexicon" from UniMorph (trust: high)
"rule" deterministic morphology rule (trust: medium)
"fallback" could not inflect; returned lemma (trust: low -> FLAG)
Public API (used by realizer_lat.py):
decline_noun(lemma, case, number) -> (form, conf)
noun_gender(lemma) -> "m"|"f"|"n"
decline_adj(lemma, case, gender, number) -> (form, conf)
conjugate(lemma, tense, mood, voice, person, number) -> (form, conf)
participle(lemma, kind, case, gender, number) -> (form, conf) # kind: prs|pfv|fut
infinitive(lemma, tense="present", voice="active") -> (form, conf)
lexicon_stats() -> dict
"""
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "lat.unimorph")
_CACHE = os.path.join(_HERE, "data", "lat_morph_cache.pkl")
_CASES = ("NOM", "GEN", "DAT", "ACC", "ABL", "VOC")
_CASE_MAP = {"nom": "NOM", "gen": "GEN", "dat": "DAT", "acc": "ACC",
"abl": "ABL", "voc": "VOC"}
_NUM = {"singular": "SG", "plural": "PL"}
_GEN = {"m": "MASC", "f": "FEM", "n": "NEUT"}
# ── UniMorph loader: noun + adjective + participle case paradigms ────────────────
def _build_cache():
nouns = {} # lemma -> {(CASE, NUM): form}
adjs = {} # lemma -> {(CASE, GEN, NUM): form}
ptcps = {} # lemma -> {(CASE, GEN, NUM): form} (from V.PTCP; keyed loosely)
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
feats = tag.split(";")
head = feats[0]
fs = set(feats)
case = next((c for c in _CASES if c in fs), None)
# handle combined case tags like GEN+DAT
if case is None:
for f in feats:
if "+" in f and any(c in f.split("+") for c in _CASES):
case = [c for c in _CASES if c in f.split("+")]
break
num = "SG" if "SG" in fs else ("PL" if "PL" in fs else None)
if case is None or num is None:
continue
cases = case if isinstance(case, list) else [case]
if head == "N":
d = nouns.setdefault(lemma, {})
for c in cases:
d.setdefault((c, num), form)
elif head == "ADJ":
# gender may be combined: MASC+FEM+NEUT, MASC+FEM
genders = []
for g in ("MASC", "FEM", "NEUT"):
if any(g == x or (g in x.split("+")) for x in feats):
genders.append(g)
if not genders:
genders = ["MASC", "FEM", "NEUT"]
d = adjs.setdefault(lemma, {})
for c in cases:
for g in genders:
d.setdefault((c, g, num), form)
data = {"nouns": nouns, "adjs": adjs, "ptcps": ptcps}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE) and os.path.exists(_UNIMORPH):
if os.path.getmtime(_CACHE) >= os.path.getmtime(_UNIMORPH):
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_NOUNS, _ADJS = _LEX["nouns"], _LEX["adjs"]
# ── noun gender inference (declension-based, curated exceptions) ─────────────────
# Real, attestable rule: gender follows declension + nominative shape, with the
# standard closed set of exceptions.
_GENDER_EXC = {
# 1st-declension masculines (people/agents)
"agricola": "m", "poēta": "m", "nauta": "m", "incola": "m", "scrība": "m",
"auriga": "m", "pīrāta": "m", "athlēta": "m",
# 2nd-declension neuters / feminines
"vīrus": "n", "vulgus": "n", "pelagus": "n", "humus": "f",
# common 3rd-declension whose gender the ending would mispredict
"rēx": "m", "dux": "m", "mīles": "m", "pater": "m", "frāter": "m",
"homō": "m", "leō": "m", "sōl": "m", "mōns": "m", "pōns": "m", "fōns": "m",
"sanguis": "m", "ōrdō": "m", "sermō": "m", "amor": "m", "dolor": "m",
"labor": "m", "timor": "m", "honor": "m", "color": "m", "pēs": "m",
"dēns": "m", "flōs": "m", "mōs": "m", "mensis": "m", "orbis": "m",
"piscis": "m", "ignis": "m", "collis": "m", "grex": "m", "prīnceps": "m",
"māter": "f", "soror": "f", "uxor": "f", "mulier": "f", "virgō": "f",
"urbs": "f", "arx": "f", "pāx": "f", "lēx": "f", "lūx": "f", "vōx": "f",
"nox": "f", "nix": "f", "vīs": "f", "salūs": "f", "virtūs": "f",
"aetās": "f", "cīvitās": "f", "lībertās": "f", "vēritās": "f", "voluptās": "f",
"nātiō": "f", "ratiō": "f", "ōrātiō": "f", "legiō": "f", "regiō": "f",
"mens": "f", "gens": "f", "ars": "f", "pars": "f", "mors": "f", "sors": "f",
"nāvis": "f", "turris": "f", "avis": "f", "vallis": "f", "classis": "f",
"corpus": "n", "tempus": "n", "opus": "n", "genus": "n", "onus": "n",
"pectus": "n", "latus": "n", "vulnus": "n", "scelus": "n", "sīdus": "n",
"caput": "n", "iter": "n", "flūmen": "n", "nōmen": "n", "carmen": "n",
"agmen": "n", "certāmen": "n", "lūmen": "n", "ōmen": "n", "cōgnōmen": "n",
"mare": "n", "animal": "n", "exemplar": "n", "rēte": "n",
# 4th-declension exceptions
"manus": "f", "domus": "f", "tribus": "f", "porticus": "f", "īdūs": "f",
"cornū": "n", "genū": "n", "gelū": "n", "verū": "n",
# 5th-declension
"diēs": "m", "merīdiēs": "m",
}
def _infer_gender(lemma):
if lemma in _GENDER_EXC:
return _GENDER_EXC[lemma]
d = _NOUNS.get(lemma)
nom = d.get(("NOM", "SG")) if d else lemma
gen = d.get(("GEN", "SG")) if d else None
nom = nom or lemma
# 5th declension: gen -eī / -ēī
if gen and (gen.endswith("") or gen.endswith("ēī")):
return "f"
# 1st declension: nom -a, gen -ae
if nom.endswith("a") and (not gen or gen.endswith("ae")):
return "f"
# 2nd declension neuter: nom -um
if nom.endswith("um"):
return "n"
# 2nd declension masc: nom -us/-er/-ir, gen -ī
if (nom.endswith("us") or nom.endswith("er") or nom.endswith("ir")) and \
(not gen or gen.endswith("ī")):
return "m"
# 4th declension: gen -ūs
if gen and gen.endswith("ūs"):
return "n" if nom.endswith("ū") else "m"
# 3rd declension neuters by common nom endings
if nom.endswith(("men", "us", "ur", "al", "ar", "e", "ma")):
# -us here is 3rd-decl neuter type (corpus) only if gen shows -oris/-eris
if nom.endswith("us") and gen and (gen.endswith("oris") or gen.endswith("eris")
or gen.endswith("uris")):
return "n"
if nom.endswith(("men", "al", "ar", "e")):
return "n"
# default 3rd-declension: masculine (most common)
return "m"
_GENDER_CACHE = {}
def noun_gender(lemma):
lemma = lemma.strip()
if lemma not in _GENDER_CACHE:
_GENDER_CACHE[lemma] = _infer_gender(lemma)
return _GENDER_CACHE[lemma]
# ── PUBLIC: noun declension ─────────────────────────────────────────────────────
def decline_noun(lemma, case, number):
lemma = lemma.strip()
C = _CASE_MAP.get(case, case.upper())
N = _NUM.get(number, number)
d = _NOUNS.get(lemma)
if d and (C, N) in d:
return d[(C, N)], "lexicon"
# abl sg often == the -e/-o form; try nom fallback
if d:
# try VOC==NOM, ACC neuter==NOM etc are already in data; last resort lemma
return lemma, "fallback"
return lemma, "fallback"
# ── PUBLIC: adjective declension ────────────────────────────────────────────────
def decline_adj(lemma, case, gender, number):
lemma = lemma.strip()
C = _CASE_MAP.get(case, case.upper())
G = _GEN.get(gender, gender.upper())
N = _NUM.get(number, number)
d = _ADJS.get(lemma)
if d and (C, G, N) in d:
return d[(C, G, N)], "lexicon"
# try other gender (some adjs listed only under MASC+FEM etc handled at load)
if d:
for altG in ("MASC", "FEM", "NEUT"):
if (C, altG, N) in d:
return d[(C, altG, N)], "lexicon"
return lemma, "fallback"
return lemma, "fallback"
# ═══════════════════════════════════════════════════════════════════════════════
# VERB RULE ENGINE (4 conjugations + curated irregulars)
# ═══════════════════════════════════════════════════════════════════════════════
# Curated principal parts for common attested verbs:
# lemma -> (conj, present_stem, perfect_stem, supine_stem)
# conj in {1,2,3,"3io",4}. Stems carry macrons (matching UniMorph orthography).
_VERBS = {
"amō": (1, "am", "amāv", "amāt"),
"laudō": (1, "laud", "laudāv", "laudāt"),
"portō": (1, "port", "portāv", "portāt"),
"vocō": (1, "voc", "vocāv", "vocāt"),
"": (1, "d", "ded", "dat"),
"spectō": (1, "spect", "spectāv", "spectāt"),
"pugnō": (1, "pugn", "pugnāv", "pugnāt"),
"labōrō": (1, "labōr", "labōrāv", "labōrāt"),
"necō": (1, "nec", "necāv", "necāt"),
"parō": (1, "par", "parāv", "parāt"),
"cōgitō": (1, "cōgit", "cōgitāv", "cōgitāt"),
"habitō": (1, "habit", "habitāv", "habitāt"),
"nārrō": (1, "nārr", "nārrāv", "nārrāt"),
"servō": (1, "serv", "servāv", "servāt"),
"superō": (1, "super", "superāv", "superāt"),
"oppugnō": (1, "oppugn", "oppugnāv", "oppugnāt"),
"ambulō": (1, "ambul", "ambulāv", "ambulāt"),
"clāmō": (1, "clām", "clāmāv", "clāmāt"),
"vulnerō": (1, "vulner", "vulnerāv", "vulnerāt"),
"aedificō": (1, "aedific", "aedificāv", "aedificāt"),
"expugnō": (1, "expugn", "expugnāv", "expugnāt"),
"dēfendō": (3, "dēfend", "dēfend", "dēfēns"),
"petō": (3, "pet", "petīv", "petīt"),
"occīdō": (3, "occīd", "occīd", "occīs"),
"interficiō": ("3io", "interfic", "interfēc", "interfect"),
"timeō": (2, "tim", "timu", None),
"iaceō": (2, "iac", "iacu", None),
"pāreō": (2, "pār", "pāru", "pārit"),
"respondeō": (2, "respond", "respond", "respōns"),
"vertō": (3, "vert", "vert", "vers"),
"ostendō": (3, "ostend", "ostend", "ostent"),
"cōnstituō": (3, "cōnstitu", "cōnstitu", "cōnstitūt"),
"cōgnōscō": (3, "cōgnōsc", "cōgnōv", "cōgnit"),
"crēdō": (3, "crēd", "crēdid", "crēdit"),
"ēdūcō": (3, "ēdūc", "ēdūx", "ēduct"),
"cōnservō": (1, "cōnserv", "cōnservāv", "cōnservāt"),
"iuvō": (1, "iuv", "iūv", "iūt"),
"dēbeō": (2, "dēb", "dēbu", "dēbit"),
"moneō": (2, "mon", "monu", "monit"),
"videō": (2, "vid", "vīd", "vīs"),
"habeō": (2, "hab", "habu", "habit"),
"teneō": (2, "ten", "tenu", "tent"),
"timeō": (2, "tim", "timu", None),
"terreō": (2, "terr", "terru", "territ"),
"dēleō": (2, "dēl", "dēlēv", "dēlēt"),
"iubeō": (2, "iub", "iuss", "iuss"),
"maneō": (2, "man", "māns", "māns"),
"moveō": (2, "mov", "mōv", "mōt"),
"doceō": (2, "doc", "docu", "doct"),
"sedeō": (2, "sed", "sēd", "sess"),
"rīdeō": (2, "rīd", "rīs", "rīs"),
"regō": (3, "reg", "rēx", "rēct"),
"dūcō": (3, "dūc", "dūx", "duct"),
"scrībō": (3, "scrīb", "scrīps", "scrīpt"),
"mittō": (3, "mitt", "mīs", "miss"),
"pōnō": (3, "pōn", "posu", "posit"),
"agō": (3, "ag", "ēg", "āct"),
"dīcō": (3, "dīc", "dīx", "dict"),
"gerō": (3, "ger", "gess", "gest"),
"vincō": (3, "vinc", "vīc", "vict"),
"petō": (3, "pet", "petīv", "petīt"),
"legō": (3, "leg", "lēg", "lēct"),
"currō": (3, "curr", "cucurr", "curs"),
"vīvō": (3, "vīv", "vīx", "vīct"),
"quaerō": (3, "quaer", "quaesīv", "quaesīt"),
"trahō": (3, "trah", "trāx", "tract"),
"claudō": (3, "claud", "claus", "claus"),
"cōgō": (3, "cōg", "coēg", "coāct"),
"relinquō": (3, "relinqu", "relīqu", "relict"),
"capiō": ("3io", "cap", "cēp", "capt"),
"faciō": ("3io", "fac", "fēc", "fact"),
"iaciō": ("3io", "iac", "iēc", "iact"),
"rapiō": ("3io", "rap", "rapu", "rapt"),
"fugiō": ("3io", "fug", "fūg", "fugit"),
"cupiō": ("3io", "cup", "cupīv", "cupīt"),
"accipiō": ("3io", "accip", "accēp", "accept"),
"audiō": (4, "aud", "audīv", "audīt"),
"veniō": (4, "ven", "vēn", "vent"),
"sciō": (4, "sc", "scīv", "scīt"),
"sentiō": (4, "sent", "sēns", "sēns"),
"mūniō": (4, "mūn", "mūnīv", "mūnīt"),
"dormiō": (4, "dorm", "dormīv", "dormīt"),
"aperiō": (4, "aper", "aperu", "apert"),
"inveniō": (4, "inven", "invēn", "invent"),
}
# ── Present-system paradigms: full ending tables per conjugation, attached to the
# bare present stem (pstem). Hardcoded from the standard grammar with correct
# macrons/vowel-lengths — deterministic and independently verifiable. Keys:
# (tense, mood, voice) -> {conj: [1sg,2sg,3sg,1pl,2pl,3pl]}
_PARADIGM = {
("present", "ind", "active"): {
1: ["ō", "ās", "at", "āmus", "ātis", "ant"],
2: ["", "ēs", "et", "ēmus", "ētis", "ent"],
3: ["ō", "is", "it", "imus", "itis", "unt"],
"3io": ["", "is", "it", "imus", "itis", "iunt"],
4: ["", "īs", "it", "īmus", "ītis", "iunt"],
},
("present", "ind", "passive"): {
1: ["or", "āris", "ātur", "āmur", "āminī", "antur"],
2: ["eor", "ēris", "ētur", "ēmur", "ēminī", "entur"],
3: ["or", "eris", "itur", "imur", "iminī", "untur"],
"3io": ["ior", "eris", "itur", "imur", "iminī", "iuntur"],
4: ["ior", "īris", "ītur", "īmur", "īminī", "iuntur"],
},
("imperfect", "ind", "active"): {
1: ["ābam", "ābās", "ābat", "ābāmus", "ābātis", "ābant"],
2: ["ēbam", "ēbās", "ēbat", "ēbāmus", "ēbātis", "ēbant"],
3: ["ēbam", "ēbās", "ēbat", "ēbāmus", "ēbātis", "ēbant"],
"3io": ["iēbam", "iēbās", "iēbat", "iēbāmus", "iēbātis", "iēbant"],
4: ["iēbam", "iēbās", "iēbat", "iēbāmus", "iēbātis", "iēbant"],
},
("imperfect", "ind", "passive"): {
1: ["ābar", "ābāris", "ābātur", "ābāmur", "ābāminī", "ābantur"],
2: ["ēbar", "ēbāris", "ēbātur", "ēbāmur", "ēbāminī", "ēbantur"],
3: ["ēbar", "ēbāris", "ēbātur", "ēbāmur", "ēbāminī", "ēbantur"],
"3io": ["iēbar", "iēbāris", "iēbātur", "iēbāmur", "iēbāminī", "iēbantur"],
4: ["iēbar", "iēbāris", "iēbātur", "iēbāmur", "iēbāminī", "iēbantur"],
},
("future", "ind", "active"): {
1: ["ābō", "ābis", "ābit", "ābimus", "ābitis", "ābunt"],
2: ["ēbō", "ēbis", "ēbit", "ēbimus", "ēbitis", "ēbunt"],
3: ["am", "ēs", "et", "ēmus", "ētis", "ent"],
"3io": ["iam", "iēs", "iet", "iēmus", "iētis", "ient"],
4: ["iam", "iēs", "iet", "iēmus", "iētis", "ient"],
},
("future", "ind", "passive"): {
1: ["ābor", "āberis", "ābitur", "ābimur", "ābiminī", "ābuntur"],
2: ["ēbor", "ēberis", "ēbitur", "ēbimur", "ēbiminī", "ēbuntur"],
3: ["ar", "ēris", "ētur", "ēmur", "ēminī", "entur"],
"3io": ["iar", "iēris", "iētur", "iēmur", "iēminī", "ientur"],
4: ["iar", "iēris", "iētur", "iēmur", "iēminī", "ientur"],
},
("present", "sbjv", "active"): {
1: ["em", "ēs", "et", "ēmus", "ētis", "ent"],
2: ["eam", "eās", "eat", "eāmus", "eātis", "eant"],
3: ["am", "ās", "at", "āmus", "ātis", "ant"],
"3io": ["iam", "iās", "iat", "iāmus", "iātis", "iant"],
4: ["iam", "iās", "iat", "iāmus", "iātis", "iant"],
},
("present", "sbjv", "passive"): {
1: ["er", "ēris", "ētur", "ēmur", "ēminī", "entur"],
2: ["ear", "eāris", "eātur", "eāmur", "eāminī", "eantur"],
3: ["ar", "āris", "ātur", "āmur", "āminī", "antur"],
"3io": ["iar", "iāris", "iātur", "iāmur", "iāminī", "iantur"],
4: ["iar", "iāris", "iātur", "iāmur", "iāminī", "iantur"],
},
("imperfect", "sbjv", "active"): {
1: ["ārem", "ārēs", "āret", "ārēmus", "ārētis", "ārent"],
2: ["ērem", "ērēs", "ēret", "ērēmus", "ērētis", "ērent"],
3: ["erem", "erēs", "eret", "erēmus", "erētis", "erent"],
"3io": ["erem", "erēs", "eret", "erēmus", "erētis", "erent"],
4: ["īrem", "īrēs", "īret", "īrēmus", "īrētis", "īrent"],
},
("imperfect", "sbjv", "passive"): {
1: ["ārer", "ārēris", "ārētur", "ārēmur", "ārēminī", "ārentur"],
2: ["ērer", "ērēris", "ērētur", "ērēmur", "ērēminī", "ērentur"],
3: ["erer", "erēris", "erētur", "erēmur", "erēminī", "erentur"],
"3io": ["erer", "erēris", "erētur", "erēmur", "erēminī", "erentur"],
4: ["īrer", "īrēris", "īrētur", "īrēmur", "īrēminī", "īrentur"],
},
}
# perfect-active endings (added to perfect stem) — same for all conjugations
_PERF_ACT = {
("perfect", "ind"): ["ī", "istī", "it", "imus", "istis", "ērunt"],
("pluperfect", "ind"): ["eram", "erās", "erat", "erāmus", "erātis", "erant"],
("futureperfect", "ind"): ["erō", "eris", "erit", "erimus", "eritis", "erint"],
("perfect", "sbjv"): ["erim", "erīs", "erit", "erīmus", "erītis", "erint"],
("pluperfect", "sbjv"):["issem", "issēs", "isset", "issēmus", "issētis", "issent"],
}
def _idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _present_system(conj, pstem, tense, mood, voice, person, number):
"""Generate a present-system form (present/imperfect/future ind & subj)."""
table = _PARADIGM.get((tense, mood, voice))
if not table or conj not in table:
return None
return pstem + table[conj][_idx(person, number)]
def _active_infinitive_stem(conj, pstem):
return {1: pstem + "ā", 2: pstem + "ē", 3: pstem + "e",
"3io": pstem + "e", 4: pstem + "ī"}[conj]
_IRREG = {
"sum": {
("present", "ind", "active"): ["sum", "es", "est", "sumus", "estis", "sunt"],
("imperfect", "ind", "active"): ["eram", "erās", "erat", "erāmus", "erātis", "erant"],
("future", "ind", "active"): ["erō", "eris", "erit", "erimus", "eritis", "erunt"],
("perfect", "ind", "active"): ["fuī", "fuistī", "fuit", "fuimus", "fuistis", "fuērunt"],
("pluperfect", "ind", "active"): ["fueram", "fuerās", "fuerat", "fuerāmus", "fuerātis", "fuerant"],
("present", "sbjv", "active"): ["sim", "sīs", "sit", "sīmus", "sītis", "sint"],
("imperfect", "sbjv", "active"): ["essem", "essēs", "esset", "essēmus", "essētis", "essent"],
},
"possum": {
("present", "ind", "active"): ["possum", "potes", "potest", "possumus", "potestis", "possunt"],
("imperfect", "ind", "active"): ["poteram", "poterās", "poterat", "poterāmus", "poterātis", "poterant"],
("future", "ind", "active"): ["poterō", "poteris", "poterit", "poterimus", "poteritis", "poterunt"],
("perfect", "ind", "active"): ["potuī", "potuistī", "potuit", "potuimus", "potuistis", "potuērunt"],
("present", "sbjv", "active"): ["possim", "possīs", "possit", "possīmus", "possītis", "possint"],
},
"": {
("present", "ind", "active"): ["", "īs", "it", "īmus", "ītis", "eunt"],
("imperfect", "ind", "active"): ["ībam", "ībās", "ībat", "ībāmus", "ībātis", "ībant"],
("future", "ind", "active"): ["ībō", "ībis", "ībit", "ībimus", "ībitis", "ībunt"],
("perfect", "ind", "active"): ["", "īstī", "iit", "iimus", "īstis", "iērunt"],
("present", "sbjv", "active"): ["eam", "eās", "eat", "eāmus", "eātis", "eant"],
},
"volō": {
("present", "ind", "active"): ["volō", "vīs", "vult", "volumus", "vultis", "volunt"],
("imperfect", "ind", "active"): ["volēbam", "volēbās", "volēbat", "volēbāmus", "volēbātis", "volēbant"],
("future", "ind", "active"): ["volam", "volēs", "volet", "volēmus", "volētis", "volent"],
("perfect", "ind", "active"): ["voluī", "voluistī", "voluit", "voluimus", "voluistis", "voluērunt"],
("present", "sbjv", "active"): ["velim", "velīs", "velit", "velīmus", "velītis", "velint"],
},
"nōlō": {
("present", "ind", "active"): ["nōlō", "nōn vīs", "nōn vult", "nōlumus", "nōn vultis", "nōlunt"],
("present", "sbjv", "active"): ["nōlim", "nōlīs", "nōlit", "nōlīmus", "nōlītis", "nōlint"],
},
"ferō": {
("present", "ind", "active"): ["ferō", "fers", "fert", "ferimus", "fertis", "ferunt"],
("imperfect", "ind", "active"): ["ferēbam", "ferēbās", "ferēbat", "ferēbāmus", "ferēbātis", "ferēbant"],
("future", "ind", "active"): ["feram", "ferēs", "feret", "ferēmus", "ferētis", "ferent"],
("perfect", "ind", "active"): ["tulī", "tulistī", "tulit", "tulimus", "tulistis", "tulērunt"],
("present", "sbjv", "active"): ["feram", "ferās", "ferat", "ferāmus", "ferātis", "ferant"],
},
}
def conjugate(lemma, tense, mood, voice="active", person="third", number="singular"):
"""Return (surface, confidence). Perfect-passive forms are periphrastic and
handled in the realizer (sum + PPP); this returns synthetic forms only."""
lemma = lemma.strip()
i = _idx(person, number)
ir = _IRREG.get(lemma)
if ir:
tbl = ir.get((tense, mood, voice)) or ir.get((tense, mood, "active"))
if tbl and tbl[i]:
return tbl[i], "rule"
v = _VERBS.get(lemma)
if not v:
v = _infer_principal_parts(lemma)
if not v:
return lemma, "fallback"
conj, pstem, perfstem, supstem = v
# imperative (present active) 2sg / 2pl
if mood == "imp":
return _imperative(conj, pstem, person, number), "rule"
# perfect-system active
if tense in ("perfect", "pluperfect", "futureperfect") and voice == "active":
if not perfstem:
return lemma, "fallback"
end = _PERF_ACT.get((tense, mood))
if end:
return perfstem + end[i], "rule"
# present-system (active + passive)
if tense in ("present", "imperfect", "future"):
form = _present_system(conj, pstem, tense, mood, voice, person, number)
if form:
return form, "rule"
return lemma, "fallback"
def _imperative(conj, pstem, person, number):
if number == "singular":
return {1: pstem + "ā", 2: pstem + "ē", 3: pstem + "e",
"3io": pstem + "e", 4: pstem + "ī"}[conj]
return {1: pstem + "āte", 2: pstem + "ēte", 3: pstem + "ite",
"3io": pstem + "ite", 4: pstem + "īte"}[conj]
def _infer_principal_parts(lemma):
"""OOV fallback: infer conjugation + stems from the 1sg-present citation form.
Perfect/supine stems are guessed regularly (often wrong for 3rd conj) and the
resulting forms are still returned as 'rule' but the realizer down-weights."""
if lemma.endswith("ō"):
base = lemma[:-1]
# can't distinguish conj from 1sg alone reliably; default by ending vowel
if base.endswith("i"):
return ("3io", base[:-1], base[:-1] + "īv", base[:-1] + "īt")
return (3, base, base + "s", base + "t")
return None
# ── PUBLIC: participles ─────────────────────────────────────────────────────────
def participle(lemma, kind, case="nom", gender="m", number="singular"):
"""kind: 'prs' (present active, -ns/-ntis), 'pfv' (perfect passive, -tus),
'fut' (future active, -tūrus). Declined as an adjective via rule endings.
Returns (form, conf)."""
v = _VERBS.get(lemma)
if not v:
return lemma, "fallback"
conj, pstem, perfstem, supstem = v
if kind == "pfv":
if not supstem:
return lemma, "fallback"
base = supstem[:-1] if supstem.endswith("t") or supstem.endswith("s") else supstem
stem = supstem # supine stem already ends in t/s: amāt- -> amātus
return _decline_us_a_um(stem, case, gender, number), "rule"
if kind == "fut":
if not supstem:
return lemma, "fallback"
return _decline_us_a_um(supstem + "ūr", case, gender, number), "rule"
if kind == "prs":
# present active participle: stem + ns (nom), stem + nt- (oblique), 3rd-decl
pv = {1: "ā", 2: "ē", 3: "ē", "3io": "", 4: ""}[conj]
ntstem = pstem + pv + "nt"
return _decline_pres_ptcp(pstem + pv, case, gender, number), "rule"
return lemma, "fallback"
def _decline_us_a_um(stem, case, gender, number):
"""Decline a -us/-a/-um adjective/participle stem (2-1-2 declension)."""
C = _CASE_MAP.get(case, case.upper())
end = {
("NOM", "m", "singular"): "us", ("NOM", "f", "singular"): "a", ("NOM", "n", "singular"): "um",
("GEN", "m", "singular"): "ī", ("GEN", "f", "singular"): "ae", ("GEN", "n", "singular"): "ī",
("DAT", "m", "singular"): "ō", ("DAT", "f", "singular"): "ae", ("DAT", "n", "singular"): "ō",
("ACC", "m", "singular"): "um", ("ACC", "f", "singular"): "am", ("ACC", "n", "singular"): "um",
("ABL", "m", "singular"): "ō", ("ABL", "f", "singular"): "ā", ("ABL", "n", "singular"): "ō",
("VOC", "m", "singular"): "e", ("VOC", "f", "singular"): "a", ("VOC", "n", "singular"): "um",
("NOM", "m", "plural"): "ī", ("NOM", "f", "plural"): "ae", ("NOM", "n", "plural"): "a",
("GEN", "m", "plural"): "ōrum", ("GEN", "f", "plural"): "ārum", ("GEN", "n", "plural"): "ōrum",
("DAT", "m", "plural"): "īs", ("DAT", "f", "plural"): "īs", ("DAT", "n", "plural"): "īs",
("ACC", "m", "plural"): "ōs", ("ACC", "f", "plural"): "ās", ("ACC", "n", "plural"): "a",
("ABL", "m", "plural"): "īs", ("ABL", "f", "plural"): "īs", ("ABL", "n", "plural"): "īs",
("VOC", "m", "plural"): "ī", ("VOC", "f", "plural"): "ae", ("VOC", "n", "plural"): "a",
}.get((C, gender, number), "us")
return stem + end
def _decline_pres_ptcp(stem, case, gender, number):
"""Present active participle (amāns, amantis) — 3rd-declension, stem+ns/nt."""
C = _CASE_MAP.get(case, case.upper())
if C == "NOM" and number == "singular":
return stem + "ns"
if C == "VOC" and number == "singular":
return stem + "ns"
base = stem + "nt"
end = {
("GEN", "singular"): "is", ("DAT", "singular"): "ī",
("ACC", "singular"): "em" if gender != "n" else "",
("ABL", "singular"): "e",
("NOM", "plural"): "ēs" if gender != "n" else "ia",
("GEN", "plural"): "ium", ("DAT", "plural"): "ibus",
("ACC", "plural"): "ēs" if gender != "n" else "ia",
("ABL", "plural"): "ibus", ("VOC", "plural"): "ēs",
}.get((C, number), "is")
if C == "ACC" and number == "singular" and gender == "n":
return stem + "ns"
return base + end
def infinitive(lemma, tense="present", voice="active"):
lemma = lemma.strip()
if lemma == "sum":
return ("esse", "rule") if tense == "present" else ("fuisse", "rule")
v = _VERBS.get(lemma)
if not v:
return lemma, "fallback"
conj, pstem, perfstem, supstem = v
if tense == "present":
if voice == "active":
return _active_infinitive_stem(conj, pstem).rstrip() + \
("re" if conj != 3 and conj != "3io" else "re"), "rule"
# passive present infinitive
base = {1: pstem + "ā", 2: pstem + "ē", 4: pstem + "ī"}.get(conj)
if base:
return base + "", "rule"
return pstem + "ī", "rule" # 3rd: regī
if tense == "perfect" and voice == "active" and perfstem:
return perfstem + "isse", "rule"
return lemma, "fallback"
def lexicon_stats():
return {
"noun_adj_source": "UniMorph Latin (github.com/unimorph/lat, CC-BY-SA 3.0)",
"verb_source": "rule-based 4-conjugation engine over curated attested "
"principal parts (UniMorph verb list is a 947-lemma sample "
"MISSING all core verbs — amō/sum/videō absent)",
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
"curated_verb_lemmas": len(_VERBS) + len(_IRREG),
"gender_inference": "declension-based (nom+gen endings) + curated exceptions",
}
if __name__ == "__main__":
import json
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
print("\n-- noun declension puella (1st, fem) --")
for c in ("nom", "gen", "dat", "acc", "abl", "voc"):
print(f" {c}: sg={decline_noun('puella', c, 'singular')[0]:10} "
f"pl={decline_noun('puella', c, 'plural')[0]}")
print("\n-- rēx (3rd, m):", [decline_noun('rēx', c, 'singular')[0] for c in ('nom','gen','dat','acc','abl')])
print("-- gender: puella=", noun_gender("puella"), "rēx=", noun_gender("rēx"),
"bellum=", noun_gender("bellum"), "corpus=", noun_gender("corpus"),
"manus=", noun_gender("manus"), "diēs=", noun_gender("diēs"))
print("\n-- conjugate videō (2nd) present ind active --")
for p in ("first", "second", "third"):
for n in ("singular", "plural"):
print(f" {p[:3]}.{n[:2]}: {conjugate('videō','present','ind','active',p,n)[0]}")
print("-- amō forms:", conjugate("amō","present","ind","active","first","singular")[0],
conjugate("amō","imperfect","ind","active","third","plural")[0],
conjugate("amō","future","ind","active","first","singular")[0],
conjugate("amō","perfect","ind","active","third","singular")[0])
print("-- sum:", [conjugate("sum","present","ind","active",p,"singular")[0] for p in ("first","second","third")])
print("-- participle amō pfv acc.f.sg:", participle("amō","pfv","acc","f","singular")[0])
print("-- infinitive amō:", infinitive("amō")[0], "| regō pass:", infinitive("regō", voice="passive")[0])
-538
View File
@@ -1,538 +0,0 @@
"""morphology_pt_full.py — production-grade Brazilian-Portuguese morphological generator.
NOT a toy. Backed by two real, broad, Wiktionary-lineage lexicons:
VERBS UniMorph Portuguese (github.com/unimorph/por, CC-BY-SA 3.0)
4,001 verb lemmas × full paradigm (283,991 finite/non-finite forms +
20,005 participle forms). Every mood/tense pt actually inflects:
indicative present / preterite (PST;PFV) / imperfect (PST;IPFV) /
pluperfect-simple (PST;PRF) / future,
conditional (futuro do pretérito),
subjunctive present / imperfect / FUTURE (PT-specific live tense),
affirmative + negative imperative,
PERSONAL infinitive (V;{p};{n};NFIN a PT-specific finite-ish form),
past participle (4 gender/number forms) + gerúndio (V.PTCP;PRS).
NOUNS + ADJECTIVES kaikki.org Portuguese (Wiktionary extract, same lineage)
81,138 noun lemmas WITH inherent gender + real (often irregular) plural
so -ão-ões / -ãos / -ães / -õos is resolved PER LEMMA by Wiktionary,
never guessed (mãomãos, pãopães, coraçãocorações).
40,252 adjective lemmas with real feminine + masc/fem plural forms.
Fallbacks (degrade, never crash, on out-of-vocabulary input):
verbs : rule generator for regular -ar/-er/-ir paradigms
nouns : gender heuristic (endings) + rule pluralization (with -ão FLAGGED)
adjs : -o/-a gender rule + rule pluralization
Confidence flag on every form:
"lexicon" straight from UniMorph/kaikki (trust: high)
"rule" deterministic rule (trust: medium)
"fallback" could not inflect; returned lemma (trust: low -> FLAG)
Public API (used by realizer_pt.py):
conjugate(lemma, mood, tense, person, number) -> (form, conf)
personal_infinitive(lemma, person, number) -> (form, conf)
participle(lemma, gender="m", number="singular") -> (form, conf)
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"
inflect_noun(lemma, number, gender=None) -> (form, conf)
inflect_adj(lemma, gender, number) -> (form, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "por.unimorph")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_pt.jsonl")
_CACHE = os.path.join(_HERE, "data", "pt_morph_cache.pkl")
# ── mood/tense pair -> UniMorph feature triple (a in tag; b in tag; c in tag) ────
_VERB_KEYMAP = {
("ind", "present"): ("IND", "PRS", None),
("ind", "preterite"): ("IND", "PST", "PFV"),
("ind", "imperfect"): ("IND", "PST", "IPFV"),
("ind", "pluperfect"): ("IND", "PST", "PRF"), # simple mais-que-perfeito
("ind", "future"): ("IND", "FUT", None),
("ind", "conditional"): ("COND", None, None),
("sbjv", "present"): ("SBJV", "PRS", None),
("sbjv", "imperfect"): ("SBJV", "PST", "IPFV"),
("sbjv", "future"): ("SBJV", "FUT", None), # PT-specific
("imp", "affirmative"): ("IMP", "POS", None),
("imp", "negative"): ("IMP", "NEG", None),
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── build the compact lexicon from UniMorph (verbs) + kaikki (nouns/adjs) ────────
def _build_verbs():
verbs = {} # (lemma, "mood|tense|person|number") -> form
pinf = {} # (lemma, "person|number") -> personal-infinitive form
part = {} # lemma -> {("m","SG"): form, ...} past participle
ger = {} # lemma -> gerúndio
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f: # past participle: falado/falada/falados/faladas
g = "m" if "MASC" in f else ("f" if "FEM" in f else "m")
num = "SG" if "SG" in f else ("PL" if "PL" in f else "SG")
part.setdefault(lemma, {})[(g, num)] = form
elif "PRS" in f: # gerúndio: falando
ger.setdefault(lemma, form)
continue
if head != "V":
continue
# personal / impersonal infinitive
if "NFIN" in f:
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person and number:
pinf[(lemma, f"{person}|{number}")] = form
continue
# finite forms
mt = None
for (mood, tense), (a, b, c) in _VERB_KEYMAP.items():
if a not in f:
continue
if b is not None and b not in f:
continue
if c is not None and c not in f:
continue
# IND;PST needs exactly PFV|IPFV|PRF — reject if the required one absent
mt = (mood, tense)
break
if mt is None:
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
verbs.setdefault((lemma, f"{mt[0]}|{mt[1]}|{person}|{number}"), form)
return verbs, pinf, part, ger
def _kaikki_gender(arg):
if not arg:
return None
a = arg.lower()
if a.startswith("f"):
return "f"
if a.startswith("m"):
return "m"
return None
def _build_nouns_adjs():
nouns = {} # lemma -> {"g","SG","PL"}
adjs = {} # lemma -> {("m","SG"),("f","SG"),("m","PL"),("f","PL")}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
pos = d.get("pos")
word = d.get("word", "")
if not word or " " in word: # skip multiword entries
continue
forms = d.get("forms", []) or []
if pos == "noun":
ht = d.get("head_templates") or []
g = None
if ht:
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
if g is None:
tags = d.get("tags") or []
if "feminine" in tags:
g = "f"
elif "masculine" in tags:
g = "m"
pl = None
for x in forms:
t = x.get("tags") or []
if "plural" in t and "alternative" not in t and "obsolete" not in t:
pl = x.get("form")
break
# first entry wins; but a later entry with a plural fills a gap
if word not in nouns:
nouns[word] = {"g": g, "SG": word, "PL": pl}
else:
cur = nouns[word]
if cur.get("g") is None and g:
cur["g"] = g
if not cur.get("PL") and pl:
cur["PL"] = pl
elif pos == "adj":
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word)
for x in forms:
t = set(x.get("tags") or [])
fm = x.get("form")
if not fm or ("alternative" in t) or ("obsolete" in t):
continue
if "comparative" in t or "superlative" in t or \
"diminutive" in t or "augmentative" in t:
continue
if "feminine" in t and "plural" in t:
d0[("f", "PL")] = fm
elif "masculine" in t and "plural" in t:
d0[("m", "PL")] = fm
elif "feminine" in t:
d0[("f", "SG")] = fm
elif "plural" in t: # invariant-gender adj (feliz -> felizes)
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
return nouns, adjs
def _build_cache():
verbs, pinf, part, ger = _build_verbs()
nouns, adjs = _build_nouns_adjs()
data = {"verbs": verbs, "pinf": pinf, "part": part, "ger": ger,
"nouns": nouns, "adjs": adjs}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
newest_src = max(os.path.getmtime(_UNIMORPH),
os.path.getmtime(_KAIKKI) if os.path.exists(_KAIKKI) else 0)
if os.path.getmtime(_CACHE) >= newest_src:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PINF, _PART, _GER, _NOUNS, _ADJS = (
_LEX["verbs"], _LEX["pinf"], _LEX["part"], _LEX["ger"],
_LEX["nouns"], _LEX["adjs"])
# ── regular-ending rule fallback (deterministic, last resort) ────────────────────
def _vclass(lemma):
return lemma[-2:] if lemma[-2:] in ("ar", "er", "ir") else None
def _stem(lemma):
return lemma[:-2]
# endings indexed [1sg,2sg,3sg,1pl,2pl,3pl]
_REG = {
("ind", "present", "ar"): ["o", "as", "a", "amos", "ais", "am"],
("ind", "present", "er"): ["o", "es", "e", "emos", "eis", "em"],
("ind", "present", "ir"): ["o", "es", "e", "imos", "is", "em"],
("ind", "preterite", "ar"): ["ei", "aste", "ou", "amos", "astes", "aram"],
("ind", "preterite", "er"): ["i", "este", "eu", "emos", "estes", "eram"],
("ind", "preterite", "ir"): ["i", "iste", "iu", "imos", "istes", "iram"],
("ind", "imperfect", "ar"): ["ava", "avas", "ava", "ávamos", "áveis", "avam"],
("ind", "imperfect", "er"): ["ia", "ias", "ia", "íamos", "íeis", "iam"],
("ind", "imperfect", "ir"): ["ia", "ias", "ia", "íamos", "íeis", "iam"],
("sbjv", "present", "ar"): ["e", "es", "e", "emos", "eis", "em"],
("sbjv", "present", "er"): ["a", "as", "a", "amos", "ais", "am"],
("sbjv", "present", "ir"): ["a", "as", "a", "amos", "ais", "am"],
("sbjv", "imperfect", "ar"): ["asse", "asses", "asse", "ássemos", "ásseis", "assem"],
("sbjv", "imperfect", "er"): ["esse", "esses", "esse", "êssemos", "êsseis", "essem"],
("sbjv", "imperfect", "ir"): ["isse", "isses", "isse", "íssemos", "ísseis", "issem"],
("sbjv", "future", "ar"): ["ar", "ares", "ar", "armos", "ardes", "arem"],
("sbjv", "future", "er"): ["er", "eres", "er", "ermos", "erdes", "erem"],
("sbjv", "future", "ir"): ["ir", "ires", "ir", "irmos", "irdes", "irem"],
}
# future & conditional attach to the FULL infinitive
_FUT = ["ei", "ás", "á", "emos", "eis", "ão"]
_COND = ["ia", "ias", "ia", "íamos", "íeis", "iam"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
st, i = _stem(lemma), _slot_idx(person, number)
if mood == "ind" and tense == "future":
return lemma + _FUT[i]
if mood == "ind" and tense == "conditional":
return lemma + _COND[i]
if mood == "imp": # affirmative tú/vocês imperative ~ subjunctive present
table = _REG.get(("sbjv", "present", vc))
if table and tense == "negative":
return st + table[i]
# affirmative 2sg = 3sg present indicative; others = subjunctive
pres = _REG.get(("ind", "present", vc))
if person == "second" and number == "singular":
return st + pres[2]
return st + table[i] if table else None
table = _REG.get((mood, tense, vc))
if table:
return st + table[i]
return None
# verified corrections to UniMorph data errors (each audited individually, not
# guessed). The three 1PL-present entries are glued-allomorph errors surfaced by a
# full-lexicon scan for a non-final "mos" in V;1;PL;IND;PRS forms (the ONLY three).
_VERB_FIX = {
("estar", "ind", "imperfect", "third", "plural"): "estavam", # was "estávam"
("estar", "ind", "present", "first", "plural"): "estamos", # was "estamosestámos"
("haver", "ind", "present", "first", "plural"): "havemos", # was "havemoshemos"
("ir", "ind", "present", "first", "plural"): "vamos", # was "vamosimos"
}
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
"""Return (surface, confidence). mood in ind|sbjv|imp; tense per _VERB_KEYMAP."""
lemma = lemma.strip().lower()
fix = _VERB_FIX.get((lemma, mood, tense, person, number))
if fix:
return fix, "lexicon"
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
if form:
# pt-BR normalization: UniMorph `por` carries the EUROPEAN spelling of
# the -ar 1pl PRETERITE (-ámos). Brazilian PT drops the accent
# (falámos->falamos, chegámos->chegamos) — 3,334/4,001 verbs affected.
if (mood == "ind" and tense == "preterite" and person == "first"
and number == "plural" and form.endswith("ámos")):
form = form[:-4] + "amos"
return form, "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r:
return r, "rule"
return lemma, "fallback"
def personal_infinitive(lemma, person, number):
"""PT personal (inflected) infinitive: para falarmos, ao chegarem."""
lemma = lemma.strip().lower()
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _PINF.get((lemma, f"{p}|{n}"))
if form:
return form, "lexicon"
# rule: infinitive + personal endings (-, -es, -, -mos, -des, -em)
end = {("first", "singular"): "", ("second", "singular"): "es",
("third", "singular"): "", ("first", "plural"): "mos",
("second", "plural"): "des", ("third", "plural"): "em"}.get((person, number), "")
return lemma + end, "rule"
# ── PUBLIC: participle + gerund ───────────────────────────────────────────────────
def participle(lemma, gender="m", number="singular"):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
d = _PART.get(lemma)
if d:
form = d.get((g, num)) or d.get(("m", "SG"))
if form:
return form, "lexicon"
if lemma.endswith("ar"):
base = lemma[:-2] + "ad"
elif lemma[-2:] in ("er", "ir"):
base = lemma[:-2] + "id"
else:
return lemma, "fallback"
suf = {"m|SG": "o", "f|SG": "a", "m|PL": "os", "f|PL": "as"}[f"{g}|{num}"]
return base + suf, "rule"
def gerund(lemma):
lemma = lemma.strip().lower()
if lemma in _GER:
return _GER[lemma], "lexicon"
if lemma.endswith("ar"):
return lemma[:-2] + "ando", "rule"
if lemma.endswith("er"):
return lemma[:-2] + "endo", "rule"
if lemma.endswith("ir"):
return lemma[:-2] + "indo", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
_FEM_SUF = ("ção", "são", "ção", "dade", "tade", "agem", "igem", "ugem", "gem",
"ez", "eza", "ice", "ície", "tude", "ude", "âncbefore")
_FEM_SUF = ("ção", "são", "dade", "tade", "agem", "gem", "eza", "ez", "ice",
"tude", "ude", "ância", "ência", "ínia")
_MASC_SUF = ("ema", "oma", "ama", "grama", "eta", "ão") # Greek -ma etc. (mostly m)
def _gender_heuristic(noun):
for suf in _FEM_SUF:
if noun.endswith(suf):
return "f"
if noun.endswith(("ema", "oma", "ama")): # problema, idioma, programa
return "m"
if noun.endswith("a") or noun.endswith("ã"):
return "f"
if noun.endswith("o") or noun.endswith(("l", "r", "z", "m", "u", "i")):
return "m"
return "m"
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g"):
return d["g"]
return _gender_heuristic(lemma)
_INVARIANT_PL_SUF = ("s",) # paroxytones ending -s are invariant (o lápis / os lápis)
def _rule_plural(noun):
"""Deterministic PT pluralization. Returns (form, ok) where ok=False flags an
ambiguous -ão that should lower confidence (the lexicon normally resolves it)."""
if not noun:
return noun, True
if noun.endswith("ão"):
return noun[:-2] + "ões", False # majority rule, but AMBIGUOUS -> flag
if noun.endswith("m"):
return noun[:-1] + "ns", True # homem->homens, jardim->jardins
if noun.endswith("al"):
return noun[:-2] + "ais", True
if noun.endswith("el"):
return noun[:-2] + "éis", True
if noun.endswith("ol"):
return noun[:-2] + "óis", True
if noun.endswith("ul"):
return noun[:-2] + "uis", True
if noun.endswith("il"):
return noun[:-2] + "is", True # stressed (funil->funis); unstressed rarer
if noun.endswith(("r", "z")):
return noun + "es", True # flor->flores, luz->luzes
if noun.endswith("s"):
# paroxytone -s (lápis, ônibus) invariant; oxytone -s (país) -> -es
return noun, True
if noun.endswith(("a", "e", "i", "o", "u", "á", "é", "í", "ó", "ú", "ã")):
return noun + "s", True
return noun + "s", True
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if number == "singular":
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
if d and d.get("PL"):
return d["PL"], "lexicon"
form, ok = _rule_plural(lemma)
return form, ("rule" if ok else "fallback")
# ── PUBLIC: adjective agreement ──────────────────────────────────────────────────
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
d = _ADJS.get(lemma)
if d:
form = d.get((g, num))
if form:
return form, "lexicon"
# build a missing plural from this gender's singular
sg = d.get((g, "SG")) or d.get(("m", "SG")) or lemma
if num == "PL":
pl, ok = _rule_plural(sg)
return pl, ("rule" if ok else "fallback")
return sg, "lexicon"
# rule fallback: -o/-a gender, then pluralize
a = lemma
if g == "f":
if a.endswith("o"):
a = a[:-1] + "a"
elif a.endswith(("ês", "or")) and not a.endswith("ior"):
a = a + "a" # português->portuguesa, trabalhador->..a
if num == "PL":
a, ok = _rule_plural(a)
return a, ("rule" if ok else "fallback")
return a, "rule"
def lexicon_stats():
return {
"verb_source": "UniMorph Portuguese (github.com/unimorph/por)",
"noun_adj_source": "kaikki.org Portuguese (Wiktionary extract)",
"license": "CC-BY-SA (Wiktionary-derived)",
"verb_forms": len(_VERBS),
"verb_lemmas": len({k[0] for k in _VERBS}),
"personal_infinitive_forms": len(_PINF),
"participle_lemmas": len(_PART),
"gerund_lemmas": len(_GER),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("falar", "ind", "present", "first", "singular", "falo"),
("comer", "ind", "present", "third", "plural", "comem"),
("partir", "ind", "present", "first", "plural", "partimos"),
("ser", "ind", "present", "third", "singular", "é"),
("ir", "ind", "preterite", "first", "singular", "fui"),
("ter", "ind", "future", "first", "singular", "terei"),
("fazer", "sbjv", "present", "first", "singular", "faça"),
("dormir", "ind", "present", "first", "singular", "durmo"),
("dar", "ind", "preterite", "third", "singular", "deu"),
("poder", "ind", "conditional", "first", "singular", "poderia"),
("fazer", "sbjv", "future", "third", "singular", "fizer"),
("estar", "ind", "present", "third", "singular", "está"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
ok += got == exp
print(f" {flag}{lemma:8} {mood}/{tense} {per[:3]}.{num[:2]} -> {got:14} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" gender: casa=", noun_gender("casa"), "problema=", noun_gender("problema"),
"mão=", noun_gender("mão"), "coração=", noun_gender("coração"),
"flor=", noun_gender("flor"))
print(" plural: mão->", inflect_noun("mão", "plural"),
"| pão->", inflect_noun("pão", "plural"),
"| animal->", inflect_noun("animal", "plural"),
"| coração->", inflect_noun("coração", "plural"))
print(" adj: bonito/f/sg->", inflect_adj("bonito", "f", "singular"),
"| feliz/m/pl->", inflect_adj("feliz", "m", "plural"),
"| português/f/sg->", inflect_adj("português", "f", "singular"))
print(" part: fazer/m/sg->", participle("fazer"), "| ger falar->", gerund("falar"))
print(" pinf falar 1pl->", personal_infinitive("falar", "first", "plural"))
-609
View File
@@ -1,609 +0,0 @@
# -*- coding: utf-8 -*-
"""morphology_ro_full.py — production-grade Romanian morphological generator.
Romanian is the BIG typological delta of the Romance family. The verb engine and
the confidence/fallback contract TRANSFER from the Italian sibling; the NOMINAL
system is genuinely new: Romanian has a SUFFIXED definite article, a preserved
NOM/ACC vs GEN/DAT case distinction, a NEUTER gender (masc-agreeing in SG,
fem-agreeing in PL), and a VOCATIVE. Those are grounded in real per-lemma data,
not guessed.
Real, Wiktionary-lineage lexical sources:
VERBS UniMorph Romanian (github.com/unimorph/ron, CC-BY-SA 3.0)
~1216 verb lemmas × paradigm, CLEAN orthography:
indicativ prezent / imperfect (PST;IPFV) / perfectul simplu (PST;PFV) /
conjunctiv prezent (SBJV;PRS, stored WITHOUT the '' particle),
participiu (V.PTCP;PST, INVARIABLE in the perfect compus),
gerunziu (V.CVB;PRS), infinitiv (NFIN), imperativ.
ro_irreg_verbs (embedded) high-frequency verbs UniMorph MISSES
(avea, vrea, da) + the auxiliary clitic paradigms the compound tenses need
(perfect-compus am/ai/a/am/ați/au, viitor voi/vei/va/vom/veți/vor,
condițional /ai/ar/am/ați/ar). Real standard forms.
NOUNS kaikki.org Romanian (Wiktionary extract, CC-BY-SA 3.0)
the FULL declension per lemma, cleanly tagged:
(nom/acc | gen/dat | vocative) × (indefinite | definite) × (sg | pl).
This is what makes the suffixed article LEXICALLY grounded (omomul,
casăcasa, băiatbăiatul, casei gen/dat, omule vocative). Inherent gender
m / f / n (NEUTER available directly) from the head template.
ADJECTIVES UniMorph Romanian ADJ
full case × gender(MASC/FEM/NEUT) × number × definiteness paradigm.
Fallbacks (degrade, never crash, on OOV): rule verb conjugation for -a/-ea/-e/-i/-î
classes, rule pluralization, rule suffixed-article by gender+ending. Every form
carries a confidence flag: "lexicon" | "rule" | "fallback".
Public API (used by realizer_ro.py):
conjugate(lemma, mood, tense, person, number) -> (form, conf)
aux(kind, person, number) -> str # perfect / future / conditional clitics
participle(lemma) -> (form, conf) # INVARIABLE
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"|"n"
definite_suffix(noun, gender, number, case) -> (form, conf) # rule engine
inflect_noun(lemma, number, gender=None, case="nomacc", definite=False) -> (form, conf)
inflect_adj(lemma, gender, number, case="nomacc", definite=False) -> (form, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "ron.unimorph")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_ro.jsonl")
_CACHE = os.path.join(_HERE, "data", "ro_morph_cache.pkl")
# ── (mood, tense) -> UniMorph feature set ─────────────────────────────────────────
_VERB_KEYMAP = {
("ind", "present"): {"IND", "PRS"},
("ind", "imperfect"): {"IND", "PST", "IPFV"},
("ind", "perfect_s"): {"IND", "PST", "PFV"}, # perfectul simplu (regional/lit.)
("sbjv", "present"): {"SBJV", "PRS"},
("imp", "affirmative"): {"POS", "IMP"},
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── high-frequency irregulars UniMorph misses + auxiliary clitic paradigms ────────
# Real standard Romanian forms (textbook paradigms).
_IRREG = {
"avea": {
"ind|present|1|SG": "am", "ind|present|2|SG": "ai", "ind|present|3|SG": "are",
"ind|present|1|PL": "avem", "ind|present|2|PL": "aveți", "ind|present|3|PL": "au",
"ind|imperfect|1|SG": "aveam", "ind|imperfect|2|SG": "aveai",
"ind|imperfect|3|SG": "avea", "ind|imperfect|1|PL": "aveam",
"ind|imperfect|2|PL": "aveați", "ind|imperfect|3|PL": "aveau",
"sbjv|present|3|SG": "aibă", "sbjv|present|3|PL": "aibă",
"sbjv|present|1|SG": "am", "sbjv|present|2|SG": "ai",
"sbjv|present|1|PL": "avem", "sbjv|present|2|PL": "aveți",
"part": "avut", "ger": "având",
},
"vrea": {
"ind|present|1|SG": "vreau", "ind|present|2|SG": "vrei", "ind|present|3|SG": "vrea",
"ind|present|1|PL": "vrem", "ind|present|2|PL": "vreți", "ind|present|3|PL": "vor",
"ind|imperfect|1|SG": "voiam", "ind|imperfect|3|SG": "voia",
"sbjv|present|3|SG": "vrea", "sbjv|present|3|PL": "vrea",
"part": "vrut", "ger": "vrând",
},
"da": {
"ind|present|1|SG": "dau", "ind|present|2|SG": "dai", "ind|present|3|SG": "",
"ind|present|1|PL": "dăm", "ind|present|2|PL": "dați", "ind|present|3|PL": "dau",
"ind|imperfect|1|SG": "dădeam", "ind|imperfect|3|SG": "dădea",
"sbjv|present|3|SG": "dea", "sbjv|present|3|PL": "dea",
"part": "dat", "ger": "dând",
},
"fi": { # a fi — present is in UniMorph but keep participle + subjunctive here
"part": "fost", "ger": "fiind",
"sbjv|present|1|SG": "fiu", "sbjv|present|2|SG": "fii", "sbjv|present|3|SG": "fie",
"sbjv|present|1|PL": "fim", "sbjv|present|2|PL": "fiți", "sbjv|present|3|PL": "fie",
"ind|imperfect|1|SG": "eram", "ind|imperfect|2|SG": "erai",
"ind|imperfect|3|SG": "era", "ind|imperfect|1|PL": "eram",
"ind|imperfect|2|PL": "erați", "ind|imperfect|3|PL": "erau",
},
}
# auxiliary clitic paradigms (person,number)->form
_AUX = {
"perfect": {("first", "singular"): "am", ("second", "singular"): "ai",
("third", "singular"): "a", ("first", "plural"): "am",
("second", "plural"): "ați", ("third", "plural"): "au"},
"future": {("first", "singular"): "voi", ("second", "singular"): "vei",
("third", "singular"): "va", ("first", "plural"): "vom",
("second", "plural"): "veți", ("third", "plural"): "vor"},
"conditional": {("first", "singular"): "", ("second", "singular"): "ai",
("third", "singular"): "ar", ("first", "plural"): "am",
("second", "plural"): "ați", ("third", "plural"): "ar"},
}
def aux(kind, person, number):
return _AUX[kind][(person, number)]
# ── build verb lexicon from UniMorph ──────────────────────────────────────────────
def _build_verbs():
verbs, part, ger = {}, {}, {}
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f:
part.setdefault(lemma, form)
continue
if head == "V.CVB":
if "PRS" in f:
ger.setdefault(lemma, form)
continue
if head != "V":
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
# conjunctiv forms in UniMorph carry a leading 'să ' — strip it
surf = form
if surf.startswith(""):
surf = surf[3:]
for (mood, tense), req in _VERB_KEYMAP.items():
if not req <= f:
continue
if tense == "imperfect" and "PFV" in f:
continue
if tense == "perfect_s" and "IPFV" in f:
continue
# keep IND;PRS out of the PRF slot (mai-mult-ca-perfect etc. ignored)
if {"IND", "PRS"} <= req and "PRF" in f:
continue
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), surf)
break
return verbs, part, ger
# ── kaikki nouns: full declension paradigm per lemma ──────────────────────────────
_EXCL = {"alternative", "archaic", "obsolete", "regional", "dialectal", "rare",
"table-tags", "inflection-template", "error-unrecognized-form",
"diminutive", "augmentative", "informal"}
def _noun_key(tagset):
if tagset & _EXCL:
return None
if "vocative" in tagset:
case = "voc"
elif "genitive" in tagset or "dative" in tagset:
case = "gendat"
elif "nominative" in tagset or "accusative" in tagset:
case = "nomacc"
else:
return None
definite = "definite" in tagset and "indefinite" not in tagset
number = "PL" if "plural" in tagset else ("SG" if "singular" in tagset else None)
if number is None:
return None
return (case, definite, number)
def _build_nouns():
nouns = {} # lemma -> {"g":..., para:{(case,def,num):form}, "PL":plain_plural}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
if d.get("pos") != "noun":
continue
word = d.get("word", "")
if not word or " " in word:
continue
ht = d.get("head_templates") or []
g = None
if ht:
a = str((ht[0].get("args") or {}).get("1") or "").lower()
if a[:1] in ("m", "f", "n"):
g = a[:1]
entry = nouns.setdefault(word, {"g": g, "para": {}, "PL": None})
if entry["g"] is None and g:
entry["g"] = g
for x in (d.get("forms") or []):
fm = x.get("form")
tg = set(x.get("tags") or [])
if not fm or fm in ("-", "#", "") or " " in fm:
continue
if tg == {"plural"} and not entry["PL"]:
entry["PL"] = fm
k = _noun_key(tg)
if k and k not in entry["para"]:
entry["para"][k] = fm
return nouns
# ── adjectives from kaikki (UniMorph ron ADJ is sparse AND mis-tagged; kaikki is
# clean: the 4-form agreement pattern bun/bună/buni/bune). Neuter maps sg->masc,
# pl->fem, so 4 forms (m/f × SG/PL) fully cover it. ────────────────────────────
def _build_adjs():
adjs = {} # lemma -> {(gender,number): form} gender in {m,f}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
if d.get("pos") != "adj":
continue
word = d.get("word", "")
if not word or " " in word:
continue
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word) # masc sg = headword
for x in (d.get("forms") or []):
fm = x.get("form")
t = set(x.get("tags") or [])
if not fm or " " in fm or fm in ("-", "#") or (t & _EXCL):
continue
if "definite" in t or "genitive" in t or "dative" in t:
continue # keep indefinite nom/acc agr set
pl = "plural" in t
fem = "feminine" in t
masc = "masculine" in t
if fem and pl:
d0.setdefault(("f", "PL"), fm)
elif masc and pl:
d0.setdefault(("m", "PL"), fm)
elif fem and not pl:
d0.setdefault(("f", "SG"), fm)
elif pl and not fem and not masc: # bare plural -> both genders
d0.setdefault(("m", "PL"), fm)
d0.setdefault(("f", "PL"), fm)
return adjs
def _build_cache():
verbs, part, ger = _build_verbs()
nouns = _build_nouns()
adjs = _build_adjs()
data = {"verbs": verbs, "part": part, "ger": ger, "nouns": nouns, "adjs": adjs}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [_UNIMORPH, _KAIKKI]
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PART, _GER, _NOUNS, _ADJS = (
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"])
# ── rule verb conjugation fallback ────────────────────────────────────────────────
def _vclass(lemma):
if lemma.endswith("a"):
return "a"
if lemma.endswith("ea"):
return "ea"
if lemma.endswith("e"):
return "e"
if lemma.endswith("i"):
return "i"
if lemma.endswith("î"):
return "î"
return None
# regular present endings by class [1sg,2sg,3sg,1pl,2pl,3pl]
_REG_PRS = {
"a": ["", "i", "ă", "ăm", "ați", "ă"], # a lucra type (simplified)
"ea": ["", "i", "e", "em", "eți", "", ],
"e": ["", "i", "e", "em", "eți", ""],
"i": ["esc", "ești", "ește", "im", "iți", "esc"], # -i type (a vorbi)
"î": ["ăsc", "ăști", "ăște", "âm", "âți", "ăsc"],
}
_SLOT = {("first", "singular"): 0, ("second", "singular"): 1, ("third", "singular"): 2,
("first", "plural"): 3, ("second", "plural"): 4, ("third", "plural"): 5}
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
i = _SLOT[(person, number)]
body = lemma[:-len(vc)]
if mood == "ind" and tense == "present":
end = _REG_PRS[vc][i]
return body + end
if mood == "ind" and tense == "imperfect":
# -a/-i/-î -> stem + a/eai...; -e/-ea -> eam. Simplified regular imperfect.
stem = body
endings = {"a": ["am", "ai", "a", "am", "ați", "au"],
"i": ["eam", "eai", "ea", "eam", "eați", "eau"],
"î": ["am", "ai", "a", "am", "ați", "au"],
"e": ["eam", "eai", "ea", "eam", "eați", "eau"],
"ea": ["eam", "eai", "ea", "eam", "eați", "eau"]}[vc]
return stem + endings[i]
return None
# ── PUBLIC verb API ───────────────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
lemma = lemma.strip().lower()
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{_NUMBER.get(number,'?')}"
ir = _IRREG.get(lemma)
if ir and key in ir:
return ir[key], "lexicon"
form = _VERBS.get((lemma, key))
if form:
return form, "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r is not None:
return r, "rule"
return lemma, "fallback"
def participle(lemma):
"""Past participle — INVARIABLE in the perfect compus (am mers, am văzut)."""
lemma = lemma.strip().lower()
ir = _IRREG.get(lemma)
if ir and "part" in ir:
return ir["part"], "lexicon"
if lemma in _PART:
return _PART[lemma], "lexicon"
vc = _vclass(lemma)
if vc == "a":
return lemma[:-1] + "at", "rule"
if vc in ("ea",):
return lemma[:-2] + "ut", "rule"
if vc == "i":
return lemma[:-1] + "it", "rule"
if vc == "î":
return lemma[:-1] + "ât", "rule"
if vc == "e":
return lemma[:-1] + "ut", "rule"
return lemma, "fallback"
def gerund(lemma):
lemma = lemma.strip().lower()
ir = _IRREG.get(lemma)
if ir and "ger" in ir:
return ir["ger"], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
vc = _vclass(lemma)
if vc in ("a", "î"):
return lemma[:-1] + "ând", "rule"
if vc in ("ea", "e", "i"):
return lemma[:-len(vc)] + "ind", "rule"
return lemma, "fallback"
# ── noun gender ───────────────────────────────────────────────────────────────────
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g") in ("m", "f", "n"):
return d["g"]
if lemma.endswith(("ă", "a", "e")):
return "f"
return "m"
# ── SUFFIXED DEFINITE ARTICLE — rule engine (fallback for OOV nouns) ───────────────
def definite_suffix(noun, gender, number, case="nomacc"):
"""Attach the enclitic definite article by gender + ending. Returns (form, conf).
This is the headline Romanian-specific engine extension."""
n = noun
g = gender
if number == "singular":
if g in ("m", "n"):
if case == "gendat":
# masc/neut gen-dat definite: -lui
if n.endswith("e"):
return n + "lui", "rule" # câine -> câinelui
if n.endswith("u"):
return n + "lui", "rule"
return n + "ului", "rule" # om -> omului
# nom/acc
if n.endswith("e"):
return n + "le", "rule" # câine -> câinele
if n.endswith("u"):
return n + "l", "rule" # codru -> codrul
if n.endswith("i"):
return n + "ul", "rule"
return n + "ul", "rule" # om -> omul
# feminine singular
if case == "gendat":
# fem gen/dat definite = plural-stem + i (casei, fetei) — needs plural;
# approximated as: -ă->-ei, -e->-ei, -a->-alei
if n.endswith("ă"):
return n[:-1] + "ei", "rule" # casă -> casei
if n.endswith("e"):
return n[:-1] + "ei", "rule" # carte -> cărții(approx cartei)
if n.endswith("a"):
return n[:-1] + "lei", "rule"
return n + "i", "rule"
# fem nom/acc
if n.endswith("ă"):
return n[:-1] + "a", "rule" # casă -> casa
if n.endswith("e"):
return n[:-1] + "ea", "rule" # carte -> cartea
if n.endswith("a"):
return n + "ua", "rule" # stea -> steaua
if n.endswith("i"):
return n + "a", "rule"
return n + "a", "rule"
# plural
if case == "gendat":
base = noun
return base + "lor", "rule" # -lor for all gen/dat pl
if g == "m":
return noun + "i", "rule" # oameni -> oamenii (+i)
return noun + "le", "rule" # case -> casele, trenuri->trenurile
# ── rule pluralization (fallback) ─────────────────────────────────────────────────
def _rule_plural(noun, gender):
if gender == "f":
if noun.endswith("ă"):
return noun[:-1] + "e"
if noun.endswith("e"):
return noun[:-1] + "i"
if noun.endswith("a"):
return noun[:-1] + "le"
return noun + "e"
if gender == "n":
return noun + "uri"
# masculine
if noun.endswith(("e",)):
return noun[:-1] + "i"
return noun + "i"
# ── PUBLIC noun inflection ────────────────────────────────────────────────────────
def inflect_noun(lemma, number, gender=None, case="nomacc", definite=False):
lemma = lemma.strip().lower()
g = gender or noun_gender(lemma)
d = _NOUNS.get(lemma)
numk = "SG" if number == "singular" else "PL"
if d:
if case == "voc":
form = d["para"].get(("voc", True, numk)) or d["para"].get(("voc", False, numk))
if form:
return form, "lexicon"
# try the exact paradigm cell from kaikki (lexically grounded)
form = d["para"].get((case, definite, numk))
if form:
return form, "lexicon"
# indefinite fallbacks from the paradigm
if not definite:
form = d["para"].get(("nomacc", False, numk))
if form:
return form, "lexicon"
if numk == "PL" and d.get("PL"):
return d["PL"], "lexicon"
if numk == "SG":
return lemma, "lexicon"
# rule path
base = lemma if number == "singular" else _rule_plural(lemma, g)
if definite:
return definite_suffix(base, g, number, case)
return base, ("rule" if d is None else "lexicon")
# ── PUBLIC adjective agreement ────────────────────────────────────────────────────
def _neuter_map(gender, number):
# neuter agrees masculine in SG, feminine in PL
if gender == "n":
return "m" if number == "singular" else "f"
return gender
def inflect_adj(lemma, gender, number, case="nomacc", definite=False):
lemma = lemma.strip().lower()
numk = "SG" if number == "singular" else "PL"
eg = _neuter_map(gender, number) # neuter -> masc(SG)/fem(PL)
d = _ADJS.get(lemma)
if d:
form = d.get((eg, numk))
if form:
return form, "lexicon"
# rule fallback: 4-form pattern bun/bună/buni/bune keyed by effective gender
a = lemma
if number == "singular":
if eg == "f":
if a.endswith("e"):
return a, "rule" # mare invariant sg
if a.endswith("u"):
return a[:-1] + "ă", "rule" # nou -> nouă
if a.endswith("ă"):
return a, "rule"
return a + "ă", "rule" # bun -> bună
return a, "rule" # masc/neut sg = lemma
# plural
if eg == "f":
if a.endswith("e"):
return a[:-1] + "i", "rule" # mare -> mari
if a.endswith("u"):
return a[:-1] + "e", "rule" # nou -> noue (approx; 'noi' irr)
if a.endswith("ă"):
return a[:-1] + "e", "rule"
return a + "e", "rule" # bun -> bune
# masc/neut(SG-only)->here masc pl -> -i
if a.endswith("e"):
return a[:-1] + "i", "rule" # mare -> mari
if a.endswith("u"):
return a[:-1] + "i", "rule"
return a + "i", "rule" # bun -> buni
def lexicon_stats():
return {
"verb_source": "UniMorph Romanian (github.com/unimorph/ron) + curated "
"irregulars (avea/vrea/da + aux clitic paradigms)",
"noun_source": "kaikki.org Romanian — full case/definite/vocative declension",
"adj_source": "UniMorph Romanian ADJ (case×gender×number×definiteness)",
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
"unimorph_verb_forms": len(_VERBS),
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
"irregular_verb_lemmas": len(_IRREG),
"participle_lemmas": len(_PART),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
print("\n── SUFFIXED DEFINITE ARTICLE (the headline delta) ──")
for n, g in [("om", "m"), ("băiat", "m"), ("casă", "f"), ("carte", "f"),
("tren", "n"), ("student", "m"), ("floare", "f")]:
sg = inflect_noun(n, "singular", g, "nomacc", True)
pl = inflect_noun(n, "plural", g, "nomacc", True)
gd = inflect_noun(n, "singular", g, "gendat", True)
vo = inflect_noun(n, "singular", g, "voc", False)
print(f" {n:8}({g}) def.sg={sg[0]:12} def.pl={pl[0]:14} "
f"gen/dat.sg={gd[0]:12} voc={vo[0]}")
print("\n── NEUTER split agreement (tren: masc SG / fem PL) ──")
print(" tren nou ->", inflect_noun("tren", "singular", "n")[0],
inflect_adj("nou", "n", "singular")[0])
print(" trenuri noi->", inflect_noun("tren", "plural", "n")[0],
inflect_adj("nou", "n", "plural")[0])
print("\n── verbs ──")
for l, m, t, p, n, in [("merge", "ind", "present", "third", "singular"),
("avea", "ind", "present", "first", "singular"),
("fi", "ind", "present", "third", "singular"),
("vorbi", "ind", "present", "third", "plural"),
("face", "sbjv", "present", "third", "singular"),
("lucra", "ind", "imperfect", "third", "singular")]:
print(f" {l:8}{m}/{t:10}{p[:3]}.{n[:2]} -> {conjugate(l,m,t,p,n)}")
print(" perfect-aux(3sg):", aux("perfect", "third", "singular"),
"| future(1sg):", aux("future", "first", "singular"),
"| cond(3sg):", aux("conditional", "third", "singular"))
print(" participle merge/vedea:", participle("merge"), participle("vedea"))
-43
View File
@@ -1,43 +0,0 @@
// multilingual_gate.el - deterministic language detect + localized-phrase test.
fn mg_det(text: String, want: String) -> String {
let got: String = ml_detect(text)
let ok: String = "MISMATCH"
if str_eq(got, want) { let ok = "ok" }
return " detect(" + got + ") want=" + want + " (" + ok + ") :: " + text + "\n"
}
fn mg_ok(text: String, want: String) -> Int {
if str_eq(ml_detect(text), want) { return 1 }
return 0
}
fn run_ml_gate() -> String {
let t1: String = "Does Neuron use SQLite for storage?"
let t2: String = "Neuron, me explica cómo la saliencia forma las geometrías."
let t3: String = "O professor não leu o livro na memória."
let t4: String = "Che cosa memorizza Neuron nella memoria?"
let rep: String = "==== ELP multilingual detect + localized phrases ====\n"
let rep = rep + mg_det(t1, "en")
let rep = rep + mg_det(t2, "es")
let rep = rep + mg_det(t3, "pt")
let rep = rep + mg_det(t4, "it")
let rep = rep + " localized decline (pt): " + ml_tr("no_memory", "pt") + "\n"
let rep = rep + " localized decline (es): " + ml_tr("no_memory", "es") + "\n"
let rep = rep + " term(saliência->en): " + ml_term("saliência", "pt") + "\n"
let rep = rep + " pred(store->pt): " + ml_translate_pred("store", "pt") + "\n"
let ok: Int = 0
if mg_ok(t1, "en") == 1 { let ok = ok + 1 }
if mg_ok(t2, "es") == 1 { let ok = ok + 1 }
if mg_ok(t3, "pt") == 1 { let ok = ok + 1 }
if mg_ok(t4, "it") == 1 { let ok = ok + 1 }
let rep = rep + "-----------------------------------------------------------------\n"
let rep = rep + "language detected correctly: " + int_to_str(ok) + "/4\n"
if ok == 4 { let rep = rep + "ML GATE: PASS\n" } else { let rep = rep + "ML GATE: FAIL\n" }
return rep
}
println(run_ml_gate())
-52
View File
@@ -1,52 +0,0 @@
// propositions_gate.el - the READ primitive over memory text (native el).
// Proves triples are recovered from free memory text and that SACRED polarity
// survives extraction (a negative memory must yield a NOT-triple).
fn pg_check(text: String, want_pol: String) -> String {
let p: [String] = prop_extract_one(text, "nd-test")
let pol: String = slots_get(p, "polarity")
let ok: String = "MISMATCH"
if str_eq(pol, want_pol) { let ok = "ok" }
return " " + prop_repr(p) + " pol=" + pol + " expected=" + want_pol + " (" + ok + ")\n"
}
fn pg_pol_ok(text: String, want_pol: String) -> Int {
let p: [String] = prop_extract_one(text, "nd-test")
if str_eq(slots_get(p, "polarity"), want_pol) { return 1 }
return 0
}
fn run_prop_gate() -> String {
let m1: String = "Neuron stores memories in SQLite."
let m2: String = "The engram does not delete a memory."
let m3: String = "Salience never drops the negation."
let m4: String = "The teacher gives the book to the children."
let rep: String = "==== ELP proposition extraction (memory text -> triples) ====\n"
let rep = rep + pg_check(m1, "aff")
let rep = rep + pg_check(m2, "neg")
let rep = rep + pg_check(m3, "neg")
let rep = rep + pg_check(m4, "aff")
// multi-sentence memory: one triple per sentence, order preserved
let doc: String = "Neuron persists learning. It does not forget the library."
let props: [String] = prop_extract(doc, "nd-doc")
let rep = rep + " --- multi-sentence doc (" + int_to_str(native_list_len(props)) + " props) ---\n"
let di: Int = 0
while di < native_list_len(props) {
let rep = rep + " " + native_list_get(props, di) + "\n"
let di = di + 1
}
let ok: Int = 0
if pg_pol_ok(m1, "aff") == 1 { let ok = ok + 1 }
if pg_pol_ok(m2, "neg") == 1 { let ok = ok + 1 }
if pg_pol_ok(m3, "neg") == 1 { let ok = ok + 1 }
if pg_pol_ok(m4, "aff") == 1 { let ok = ok + 1 }
let rep = rep + "-----------------------------------------------------------------\n"
let rep = rep + "SACRED polarity correct on extraction: " + int_to_str(ok) + "/4\n"
if ok == 4 { let rep = rep + "PROP GATE: PASS\n" } else { let rep = rep + "PROP GATE: FAIL\n" }
return rep
}
println(run_prop_gate())
BIN
View File
Binary file not shown.
+105 -254
View File
@@ -10,9 +10,6 @@ el_val_t query_param(el_val_t path, el_val_t key);
el_val_t query_int(el_val_t path, el_val_t key, el_val_t default_val);
el_val_t extract_id(el_val_t path, el_val_t prefix);
el_val_t route_stats(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_act_stats(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_text_health(el_val_t method, el_val_t path, el_val_t body);
el_val_t persist_canonical(void);
el_val_t route_create_node(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_get_node(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_scan_nodes(el_val_t method, el_val_t path, el_val_t body);
@@ -20,29 +17,21 @@ el_val_t route_scan_edges(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_search(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_activate(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_create_edge(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_create_edges_batch(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_neighbors(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_strengthen(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_forget(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_create_ise(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_sync(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_save(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_load(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_health(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_embed_backfill(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_sync(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_load_merge(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_emit_ise(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_capture_knowledge(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_similarity(el_val_t method, el_val_t path, el_val_t body);
el_val_t check_auth_ok(el_val_t method, el_val_t body);
el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body);
el_val_t bind_raw;
el_val_t bind_str;
el_val_t port;
el_val_t data_dir_raw;
el_val_t data_dir;
el_val_t snapshot_path;
el_val_t boot_snap;
el_val_t parse_port(el_val_t bind) {
el_val_t colon = str_index_of(bind, EL_STR(":"));
@@ -121,40 +110,17 @@ el_val_t route_stats(el_val_t method, el_val_t path, el_val_t body) {
return 0;
}
el_val_t route_act_stats(el_val_t method, el_val_t path, el_val_t body) {
return engram_act_stats_json();
return 0;
}
el_val_t route_text_health(el_val_t method, el_val_t path, el_val_t body) {
return engram_text_health_json();
return 0;
}
el_val_t persist_canonical(void) {
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_1 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_1 = (EL_STR("/tmp/engram")); } else { _if_result_1 = (dir_raw); } _if_result_1; });
return engram_save(el_str_concat(dir, EL_STR("/snapshot.json")));
return 0;
}
el_val_t route_create_node(el_val_t method, el_val_t path, el_val_t body) {
el_val_t content = json_get_string(body, EL_STR("content"));
el_val_t nt_raw = json_get_string(body, EL_STR("node_type"));
el_val_t node_type = ({ el_val_t _if_result_2 = 0; if (str_eq(nt_raw, EL_STR(""))) { _if_result_2 = (EL_STR("Memory")); } else { _if_result_2 = (nt_raw); } _if_result_2; });
el_val_t sal_present = json_get_raw(body, EL_STR("salience"));
el_val_t salience = ({ el_val_t _if_result_3 = 0; if (str_eq(sal_present, EL_STR(""))) { _if_result_3 = (el_from_float(0.5)); } else { _if_result_3 = (json_get_float(body, EL_STR("salience"))); } _if_result_3; });
el_val_t label_raw = json_get_string(body, EL_STR("label"));
el_val_t label = ({ el_val_t _if_result_4 = 0; if (str_eq(label_raw, EL_STR(""))) { _if_result_4 = (content); } else { _if_result_4 = (label_raw); } _if_result_4; });
el_val_t imp_present = json_get_raw(body, EL_STR("importance"));
el_val_t importance = ({ el_val_t _if_result_5 = 0; if (str_eq(imp_present, EL_STR(""))) { _if_result_5 = (el_from_float(0.5)); } else { _if_result_5 = (json_get_float(body, EL_STR("importance"))); } _if_result_5; });
el_val_t conf_present = json_get_raw(body, EL_STR("confidence"));
el_val_t confidence = ({ el_val_t _if_result_6 = 0; if (str_eq(conf_present, EL_STR(""))) { _if_result_6 = (el_from_float(1.0)); } else { _if_result_6 = (json_get_float(body, EL_STR("confidence"))); } _if_result_6; });
el_val_t tier_raw = json_get_string(body, EL_STR("tier"));
el_val_t tier = ({ el_val_t _if_result_7 = 0; if (str_eq(tier_raw, EL_STR(""))) { _if_result_7 = (EL_STR("Working")); } else { _if_result_7 = (tier_raw); } _if_result_7; });
el_val_t tags = json_get_string(body, EL_STR("tags"));
el_val_t id = engram_node_full(content, node_type, label, salience, importance, confidence, tier, tags);
el_val_t saved = persist_canonical();
el_val_t node_type = json_get_string(body, EL_STR("node_type"));
if (str_eq(node_type, EL_STR(""))) {
node_type = EL_STR("Memory");
}
el_val_t salience = json_get_float(body, EL_STR("salience"));
if (salience == el_from_float(0.0)) {
salience = el_from_float(0.5);
}
el_val_t id = engram_node(content, node_type, salience);
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"id\":\""), id), EL_STR("\",\"content\":\"")), content), EL_STR("\",\"node_type\":\"")), node_type), EL_STR("\"}"));
return 0;
}
@@ -180,9 +146,11 @@ el_val_t route_scan_nodes(el_val_t method, el_val_t path, el_val_t body) {
}
el_val_t route_scan_edges(el_val_t method, el_val_t path, el_val_t body) {
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_8 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_8 = (EL_STR("/tmp/engram")); } else { _if_result_8 = (dir_raw); } _if_result_8; });
el_val_t snap_path = el_str_concat(dir, EL_STR("/.scan-export.json"));
el_val_t dir = env(EL_STR("ENGRAM_DATA_DIR"));
if (str_eq(dir, EL_STR(""))) {
dir = EL_STR("/tmp/engram");
}
el_val_t snap_path = el_str_concat(dir, EL_STR("/snapshot.json"));
engram_save(snap_path);
el_val_t snap = fs_read(snap_path);
if (str_eq(snap, EL_STR(""))) {
@@ -197,22 +165,36 @@ el_val_t route_scan_edges(el_val_t method, el_val_t path, el_val_t body) {
}
el_val_t route_search(el_val_t method, el_val_t path, el_val_t body) {
el_val_t q = ({ el_val_t _if_result_9 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_9 = (query_param(path, EL_STR("q"))); } else { _if_result_9 = (json_get_string(body, EL_STR("query"))); } _if_result_9; });
el_val_t lim_url = query_int(path, EL_STR("limit"), 0);
el_val_t lim_body = json_get_int(body, EL_STR("limit"));
el_val_t lim_either = ({ el_val_t _if_result_10 = 0; if ((lim_url > 0)) { _if_result_10 = (lim_url); } else { _if_result_10 = (lim_body); } _if_result_10; });
el_val_t limit = ({ el_val_t _if_result_11 = 0; if ((lim_either > 0)) { _if_result_11 = (lim_either); } else { _if_result_11 = (20); } _if_result_11; });
el_val_t q = EL_STR("");
if (str_eq(method, EL_STR("GET"))) {
q = query_param(path, EL_STR("q"));
} else {
q = json_get_string(body, EL_STR("query"));
}
el_val_t limit = query_int(path, EL_STR("limit"), 20);
if (limit == 0) {
limit = json_get_int(body, EL_STR("limit"));
}
if (limit == 0) {
limit = 20;
}
return engram_search_json(q, limit);
return 0;
}
el_val_t route_activate(el_val_t method, el_val_t path, el_val_t body) {
el_val_t q = ({ el_val_t _if_result_12 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_12 = (query_param(path, EL_STR("q"))); } else { _if_result_12 = (json_get_string(body, EL_STR("query"))); } _if_result_12; });
if (str_eq(q, EL_STR(""))) {
return err_json(EL_STR("missing query"));
el_val_t q = EL_STR("");
el_val_t depth = 3;
if (str_eq(method, EL_STR("GET"))) {
q = query_param(path, EL_STR("q"));
depth = query_int(path, EL_STR("depth"), 3);
} else {
q = json_get_string(body, EL_STR("query"));
el_val_t bd = json_get_int(body, EL_STR("depth"));
if (bd > 0) {
depth = bd;
}
}
el_val_t d_raw = ({ el_val_t _if_result_13 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_13 = (query_int(path, EL_STR("depth"), 3)); } else { _if_result_13 = (json_get_int(body, EL_STR("depth"))); } _if_result_13; });
el_val_t depth = ({ el_val_t _if_result_14 = 0; if ((d_raw > 0)) { _if_result_14 = (d_raw); } else { _if_result_14 = (3); } _if_result_14; });
return el_str_concat(el_str_concat(EL_STR("{\"results\":"), engram_activate_json(q, depth)), EL_STR("}"));
return 0;
}
@@ -220,51 +202,19 @@ el_val_t route_activate(el_val_t method, el_val_t path, el_val_t body) {
el_val_t route_create_edge(el_val_t method, el_val_t path, el_val_t body) {
el_val_t from_id = json_get_string(body, EL_STR("from_id"));
el_val_t to_id = json_get_string(body, EL_STR("to_id"));
el_val_t rel_raw = json_get_string(body, EL_STR("relation"));
el_val_t relation = ({ el_val_t _if_result_15 = 0; if (str_eq(rel_raw, EL_STR(""))) { _if_result_15 = (EL_STR("associates")); } else { _if_result_15 = (rel_raw); } _if_result_15; });
el_val_t w_present = json_get_raw(body, EL_STR("weight"));
el_val_t weight = ({ el_val_t _if_result_16 = 0; if (str_eq(w_present, EL_STR(""))) { _if_result_16 = (el_from_float(0.5)); } else { _if_result_16 = (json_get_float(body, EL_STR("weight"))); } _if_result_16; });
el_val_t relation = json_get_string(body, EL_STR("relation"));
if (str_eq(relation, EL_STR(""))) {
relation = EL_STR("associates");
}
el_val_t weight = json_get_float(body, EL_STR("weight"));
if (weight == el_from_float(0.0)) {
weight = el_from_float(0.5);
}
engram_connect(from_id, to_id, weight, relation);
el_val_t saved = persist_canonical();
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"from_id\":\""), from_id), EL_STR("\",\"to_id\":\"")), to_id), EL_STR("\",\"relation\":\"")), relation), EL_STR("\"}"));
return 0;
}
el_val_t route_create_edges_batch(el_val_t method, el_val_t path, el_val_t body) {
el_val_t arr = json_get_raw(body, EL_STR("edges"));
if (str_eq(arr, EL_STR(""))) {
return err_json(EL_STR("missing edges array"));
}
el_val_t n = json_array_len(arr);
if (n == 0) {
return EL_STR("{\"ok\":true,\"accepted\":0,\"skipped\":0}");
}
el_val_t i = 0;
el_val_t accepted = 0;
el_val_t skipped = 0;
while (i < n) {
el_val_t item = json_array_get(arr, i);
el_val_t from_id = json_get_string(item, EL_STR("from_id"));
el_val_t to_id = json_get_string(item, EL_STR("to_id"));
if (str_eq(from_id, EL_STR("")) || str_eq(to_id, EL_STR(""))) {
skipped = (skipped + 1);
} else {
el_val_t rel_raw = json_get_string(item, EL_STR("relation"));
el_val_t relation = ({ el_val_t _if_result_17 = 0; if (str_eq(rel_raw, EL_STR(""))) { _if_result_17 = (EL_STR("associates")); } else { _if_result_17 = (rel_raw); } _if_result_17; });
el_val_t w_present = json_get_raw(item, EL_STR("weight"));
el_val_t weight = ({ el_val_t _if_result_18 = 0; if (str_eq(w_present, EL_STR(""))) { _if_result_18 = (el_from_float(0.5)); } else { _if_result_18 = (json_get_float(item, EL_STR("weight"))); } _if_result_18; });
engram_connect(from_id, to_id, weight, relation);
accepted = (accepted + 1);
}
i = (i + 1);
}
if (accepted > 0) {
el_val_t saved = persist_canonical();
}
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"accepted\":"), int_to_str(accepted)), EL_STR(",\"skipped\":")), int_to_str(skipped)), EL_STR("}"));
return 0;
}
el_val_t route_neighbors(el_val_t method, el_val_t path, el_val_t body) {
el_val_t id = extract_id(path, EL_STR("/api/neighbors/"));
if (str_eq(id, EL_STR(""))) {
@@ -281,7 +231,6 @@ el_val_t route_strengthen(el_val_t method, el_val_t path, el_val_t body) {
return err_json(EL_STR("missing node_id"));
}
engram_strengthen(id);
el_val_t saved = persist_canonical();
return ok_json();
return 0;
}
@@ -292,83 +241,11 @@ el_val_t route_forget(el_val_t method, el_val_t path, el_val_t body) {
return err_json(EL_STR("missing id"));
}
engram_forget(id);
el_val_t saved = persist_canonical();
return ok_json();
return 0;
}
el_val_t route_save(el_val_t method, el_val_t path, el_val_t body) {
el_val_t p_raw = json_get_string(body, EL_STR("path"));
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_19 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_19 = (EL_STR("/tmp/engram")); } else { _if_result_19 = (dir_raw); } _if_result_19; });
el_val_t p = ({ el_val_t _if_result_20 = 0; if (str_eq(p_raw, EL_STR(""))) { _if_result_20 = (el_str_concat(dir, EL_STR("/snapshot.json"))); } else { _if_result_20 = (p_raw); } _if_result_20; });
el_val_t sv = engram_save(p);
el_val_t sv_ok = ({ el_val_t _if_result_21 = 0; if ((sv == 0)) { _if_result_21 = (EL_STR("false")); } else { _if_result_21 = (EL_STR("true")); } _if_result_21; });
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":"), sv_ok), EL_STR(",\"path\":\"")), p), EL_STR("\",\"node_count\":")), int_to_str(engram_node_count())), EL_STR(",\"edge_count\":")), int_to_str(engram_edge_count())), EL_STR("}"));
return 0;
}
el_val_t route_load(el_val_t method, el_val_t path, el_val_t body) {
el_val_t p_raw = json_get_string(body, EL_STR("path"));
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_22 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_22 = (EL_STR("/tmp/engram")); } else { _if_result_22 = (dir_raw); } _if_result_22; });
el_val_t p = ({ el_val_t _if_result_23 = 0; if (str_eq(p_raw, EL_STR(""))) { _if_result_23 = (el_str_concat(dir, EL_STR("/snapshot.json"))); } else { _if_result_23 = (p_raw); } _if_result_23; });
el_val_t ld = engram_load(p);
el_val_t ld_ok = ({ el_val_t _if_result_24 = 0; if ((ld == 0)) { _if_result_24 = (EL_STR("false")); } else { _if_result_24 = (EL_STR("true")); } _if_result_24; });
el_val_t nc_after = engram_node_count();
el_val_t hollow = ({ el_val_t _if_result_25 = 0; if ((nc_after == 0)) { _if_result_25 = (EL_STR("true")); } else { _if_result_25 = (EL_STR("false")); } _if_result_25; });
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":"), ld_ok), EL_STR(",\"path\":\"")), p), EL_STR("\",\"node_count\":")), int_to_str(nc_after)), EL_STR(",\"edge_count\":")), int_to_str(engram_edge_count())), EL_STR(",\"hollow\":")), hollow), EL_STR("}"));
return 0;
}
el_val_t route_health(el_val_t method, el_val_t path, el_val_t body) {
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"status\":\"ok\",\"engine\":\"engram-runtime-native\",\"node_count\":"), int_to_str(engram_node_count())), EL_STR(",\"edge_count\":")), int_to_str(engram_edge_count())), EL_STR("}"));
return 0;
}
el_val_t route_embed_backfill(el_val_t method, el_val_t path, el_val_t body) {
el_val_t n = query_int(path, EL_STR("n"), 32);
el_val_t result = engram_embed_backfill(n);
el_val_t done = json_get_float(result, EL_STR("embedded"));
if (done > el_from_float(0.0)) {
el_val_t saved = persist_canonical();
}
return result;
return 0;
}
el_val_t route_sync(el_val_t method, el_val_t path, el_val_t body) {
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_26 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_26 = (EL_STR("/tmp/engram")); } else { _if_result_26 = (dir_raw); } _if_result_26; });
el_val_t snap_path = el_str_concat(dir, EL_STR("/.sync-export.json"));
engram_save(snap_path);
el_val_t snap = fs_read(snap_path);
if (str_eq(snap, EL_STR(""))) {
return err_json(EL_STR("sync export failed: snapshot unreadable"));
}
return snap;
return 0;
}
el_val_t route_load_merge(el_val_t method, el_val_t path, el_val_t body) {
el_val_t p = json_get_string(body, EL_STR("path"));
if (str_eq(p, EL_STR(""))) {
return err_json(EL_STR("path is required"));
}
if (str_eq(fs_read(p), EL_STR(""))) {
return err_json(EL_STR("file missing or empty"));
}
el_val_t before_n = engram_node_count();
el_val_t before_e = engram_edge_count();
engram_load_merge(p);
el_val_t added_n = (engram_node_count() - before_n);
el_val_t added_e = (engram_edge_count() - before_e);
el_val_t saved = persist_canonical();
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"nodes_added\":"), int_to_str(added_n)), EL_STR(",\"edges_added\":")), int_to_str(added_e)), EL_STR(",\"node_count\":")), int_to_str(engram_node_count())), EL_STR("}"));
return 0;
}
el_val_t route_emit_ise(el_val_t method, el_val_t path, el_val_t body) {
el_val_t route_create_ise(el_val_t method, el_val_t path, el_val_t body) {
el_val_t content = json_get_string(body, EL_STR("content"));
if (str_eq(content, EL_STR(""))) {
return err_json(EL_STR("missing content"));
@@ -377,55 +254,55 @@ el_val_t route_emit_ise(el_val_t method, el_val_t path, el_val_t body) {
el_val_t imp = el_from_float(0.3);
el_val_t conf = el_from_float(0.8);
el_val_t id = engram_node_full(content, EL_STR("InternalStateEvent"), EL_STR("state-event"), sal, imp, conf, EL_STR("Episodic"), EL_STR("[\"internal-state\",\"InternalStateEvent\"]"));
el_val_t ret_raw = env(EL_STR("ENGRAM_ISE_RETENTION_MS"));
el_val_t ret_ms = ({ el_val_t _if_result_27 = 0; if (str_eq(ret_raw, EL_STR(""))) { _if_result_27 = (172800000); } else { _if_result_27 = (str_to_int(ret_raw)); } _if_result_27; });
el_val_t pruned = engram_prune_telemetry(ret_ms);
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"id\":\""), id), EL_STR("\",\"pruned\":")), int_to_str(pruned)), EL_STR("}"));
return 0;
}
el_val_t route_capture_knowledge(el_val_t method, el_val_t path, el_val_t body) {
el_val_t content = json_get_string(body, EL_STR("content"));
if (str_eq(content, EL_STR(""))) {
return err_json(EL_STR("missing content"));
}
el_val_t title = json_get_string(body, EL_STR("title"));
el_val_t label = ({ el_val_t _if_result_28 = 0; if (str_eq(title, EL_STR(""))) { _if_result_28 = (str_slice(content, 0, 60)); } else { _if_result_28 = (title); } _if_result_28; });
el_val_t category_raw = json_get_string(body, EL_STR("category"));
el_val_t category = ({ el_val_t _if_result_29 = 0; if (str_eq(category_raw, EL_STR(""))) { _if_result_29 = (EL_STR("other")); } else { _if_result_29 = (category_raw); } _if_result_29; });
el_val_t ktier_raw = json_get_string(body, EL_STR("tier"));
el_val_t ktier = ({ el_val_t _if_result_30 = 0; if (str_eq(ktier_raw, EL_STR(""))) { _if_result_30 = (EL_STR("note")); } else { _if_result_30 = (ktier_raw); } _if_result_30; });
el_val_t project = json_get_string(body, EL_STR("project"));
el_val_t tags_raw = json_get_raw(body, EL_STR("tags"));
el_val_t tags_base = ({ el_val_t _if_result_31 = 0; if (str_eq(tags_raw, EL_STR(""))) { _if_result_31 = (EL_STR("[]")); } else { _if_result_31 = (tags_raw); } _if_result_31; });
el_val_t base_len = str_len(tags_base);
el_val_t head = str_slice(tags_base, 0, (base_len - 1));
el_val_t sep = ({ el_val_t _if_result_32 = 0; if (str_eq(head, EL_STR("["))) { _if_result_32 = (EL_STR("")); } else { _if_result_32 = (EL_STR(",")); } _if_result_32; });
el_val_t safe_cat = str_replace(category, EL_STR("\""), EL_STR("'"));
el_val_t safe_tier = str_replace(ktier, EL_STR("\""), EL_STR("'"));
el_val_t safe_proj = str_replace(project, EL_STR("\""), EL_STR("'"));
el_val_t proj_tag = ({ el_val_t _if_result_33 = 0; if (str_eq(safe_proj, EL_STR(""))) { _if_result_33 = (EL_STR("")); } else { _if_result_33 = (el_str_concat(el_str_concat(EL_STR(",\"project:"), safe_proj), EL_STR("\""))); } _if_result_33; });
el_val_t tags = el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(head, sep), EL_STR("\"category:")), safe_cat), EL_STR("\",\"tier:")), safe_tier), EL_STR("\"")), proj_tag), EL_STR("]"));
el_val_t sal = el_from_float(0.5);
el_val_t imp = el_from_float(0.5);
el_val_t conf = el_from_float(0.9);
el_val_t id = engram_node_full(content, EL_STR("Knowledge"), label, sal, imp, conf, EL_STR("Semantic"), tags);
el_val_t saved = persist_canonical();
return el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"id\":\""), id), EL_STR("\"}"));
return 0;
}
el_val_t route_similarity(el_val_t method, el_val_t path, el_val_t body) {
el_val_t a = query_param(path, EL_STR("a"));
el_val_t b = query_param(path, EL_STR("b"));
if (str_eq(a, EL_STR(""))) {
return err_json(EL_STR("missing a"));
el_val_t route_sync(el_val_t method, el_val_t path, el_val_t body) {
el_val_t dir = env(EL_STR("ENGRAM_DATA_DIR"));
if (str_eq(dir, EL_STR(""))) {
dir = EL_STR("/tmp/engram");
}
if (str_eq(b, EL_STR(""))) {
return err_json(EL_STR("missing b"));
el_val_t snap_path = el_str_concat(dir, EL_STR("/sync-export.json"));
engram_save(snap_path);
el_val_t snap = fs_read(snap_path);
if (str_eq(snap, EL_STR(""))) {
return EL_STR("{\"nodes\":[],\"edges\":[]}");
}
el_val_t sim = engram_cosine_sim(a, b);
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"a\":\""), a), EL_STR("\",\"b\":\"")), b), EL_STR("\",\"cosine\":")), float_to_str(sim)), EL_STR("}"));
return snap;
return 0;
}
el_val_t route_save(el_val_t method, el_val_t path, el_val_t body) {
el_val_t p = json_get_string(body, EL_STR("path"));
if (str_eq(p, EL_STR(""))) {
el_val_t dir = env(EL_STR("ENGRAM_DATA_DIR"));
if (str_eq(dir, EL_STR(""))) {
dir = EL_STR("/tmp/engram");
}
p = el_str_concat(dir, EL_STR("/snapshot.json"));
}
engram_save(p);
return el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"path\":\""), p), EL_STR("\"}"));
return 0;
}
el_val_t route_load(el_val_t method, el_val_t path, el_val_t body) {
el_val_t p = json_get_string(body, EL_STR("path"));
if (str_eq(p, EL_STR(""))) {
el_val_t dir = env(EL_STR("ENGRAM_DATA_DIR"));
if (str_eq(dir, EL_STR(""))) {
dir = EL_STR("/tmp/engram");
}
p = el_str_concat(dir, EL_STR("/snapshot.json"));
}
engram_load(p);
return ok_json();
return 0;
}
el_val_t route_health(el_val_t method, el_val_t path, el_val_t body) {
return EL_STR("{\"status\":\"ok\",\"engine\":\"engram-runtime-native\"}");
return 0;
}
@@ -452,24 +329,15 @@ el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body) {
return route_health(method, path, body);
}
}
if (str_eq(method, EL_STR("POST")) && str_eq(clean, EL_STR("/api/neuron/state-events"))) {
return route_emit_ise(method, path, body);
if (str_eq(method, EL_STR("POST")) && str_starts_with(clean, EL_STR("/api/neuron/state-events"))) {
return route_create_ise(method, path, body);
}
if (!check_auth_ok(method, body)) {
return err_json(EL_STR("unauthorized"));
}
if (str_eq(method, EL_STR("POST")) && str_eq(clean, EL_STR("/api/neuron/knowledge/capture"))) {
return route_capture_knowledge(method, path, body);
}
if (str_eq(method, EL_STR("GET")) && (str_eq(clean, EL_STR("/api/stats")) || str_eq(clean, EL_STR("/stats")))) {
return route_stats(method, path, body);
}
if (str_eq(method, EL_STR("GET")) && (str_eq(clean, EL_STR("/api/act-stats")) || str_eq(clean, EL_STR("/act-stats")))) {
return route_act_stats(method, path, body);
}
if (str_eq(method, EL_STR("GET")) && (str_eq(clean, EL_STR("/api/text-health")) || str_eq(clean, EL_STR("/text-health")))) {
return route_text_health(method, path, body);
}
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/nodes")) || str_eq(clean, EL_STR("/nodes")))) {
return route_create_node(method, path, body);
}
@@ -488,9 +356,6 @@ el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body) {
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/edges")) || str_eq(clean, EL_STR("/edges")))) {
return route_create_edge(method, path, body);
}
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/edges/batch")) || str_eq(clean, EL_STR("/edges/batch")))) {
return route_create_edges_batch(method, path, body);
}
if (str_eq(method, EL_STR("GET")) && str_starts_with(clean, EL_STR("/api/neighbors/"))) {
return route_neighbors(method, path, body);
}
@@ -509,46 +374,32 @@ el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body) {
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/strengthen")) || str_eq(clean, EL_STR("/strengthen")))) {
return route_strengthen(method, path, body);
}
if (str_eq(method, EL_STR("GET")) && (str_eq(clean, EL_STR("/api/sync")) || str_eq(clean, EL_STR("/sync")))) {
return route_sync(method, path, body);
}
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/save")) || str_eq(clean, EL_STR("/save")))) {
return route_save(method, path, body);
}
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/load")) || str_eq(clean, EL_STR("/load")))) {
return route_load(method, path, body);
}
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/load-merge")) || str_eq(clean, EL_STR("/load-merge")))) {
return route_load_merge(method, path, body);
}
if (str_eq(method, EL_STR("GET")) && str_eq(clean, EL_STR("/api/sync"))) {
return route_sync(method, path, body);
}
if (str_eq(clean, EL_STR("/api/embed-backfill"))) {
return route_embed_backfill(method, path, body);
}
if (str_eq(method, EL_STR("GET")) && str_starts_with(clean, EL_STR("/api/similarity"))) {
return route_similarity(method, path, body);
}
return el_str_concat(el_str_concat(EL_STR("{\"error\":\"not found\",\"path\":\""), clean), EL_STR("\"}"));
return 0;
}
int main(int _argc, char** _argv) {
el_runtime_init_args(_argc, _argv);
bind_raw = env(EL_STR("ENGRAM_BIND"));
bind_str = ({ el_val_t _if_result_34 = 0; if (str_eq(bind_raw, EL_STR(""))) { _if_result_34 = (EL_STR(":8742")); } else { _if_result_34 = (bind_raw); } _if_result_34; });
bind_str = env(EL_STR("ENGRAM_BIND"));
if (str_eq(bind_str, EL_STR(""))) {
bind_str = EL_STR(":8742");
}
port = parse_port(bind_str);
data_dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
data_dir = ({ el_val_t _if_result_35 = 0; if (str_eq(data_dir_raw, EL_STR(""))) { _if_result_35 = (EL_STR("/tmp/engram")); } else { _if_result_35 = (data_dir_raw); } _if_result_35; });
data_dir = env(EL_STR("ENGRAM_DATA_DIR"));
if (str_eq(data_dir, EL_STR(""))) {
data_dir = EL_STR("/tmp/engram");
}
snapshot_path = el_str_concat(data_dir, EL_STR("/snapshot.json"));
engram_load(snapshot_path);
boot_snap = fs_read(snapshot_path);
if (!str_eq(boot_snap, EL_STR(""))) {
if (engram_node_count() == 0) {
println(EL_STR("[engram] WARNING: snapshot.json is non-empty but load produced 0 nodes \xe2\x80\x94 preserving copy at snapshot.failed-load.json"));
fs_write(el_str_concat(data_dir, EL_STR("/snapshot.failed-load.json")), boot_snap);
} else {
fs_write(el_str_concat(data_dir, EL_STR("/snapshot.boot-backup.json")), boot_snap);
}
}
println(EL_STR("[engram] runtime-native graph engine"));
println(el_str_concat(EL_STR("[engram] data_dir="), data_dir));
println(el_str_concat(EL_STR("[engram] node_count="), int_to_str(engram_node_count())));
+69 -426
View File
@@ -50,8 +50,12 @@ fn query_param(path: String, key: String) -> String {
if pos < 0 { return "" }
let after: String = str_slice(qs, pos + str_len(needle), str_len(qs))
let amp: Int = str_index_of(after, "&")
if amp < 0 { return after }
str_slice(after, 0, amp)
// SPEC-SEARCH-UPGRADE 2026-07-14: URL-decode the extracted value (%XX and
// '+' were previously passed through literally, so an encoded multi-word
// query arrived as junk tokens pre-existing GET-path defect, masked
// until search could actually rank multi-word queries).
if amp < 0 { return url_decode(after) }
url_decode(str_slice(after, 0, amp))
}
fn query_int(path: String, key: String, default_val: Int) -> Int {
@@ -76,112 +80,13 @@ fn route_stats(method: String, path: String, body: String) -> String {
engram_stats_json()
}
// route_act_stats GET /api/act-stats
// (2026-08-04 self-review) engram_act_stats_json() has existed since the
// 2026-07-27 review but was reachable ONLY through the soul daemon's heartbeat
// binding. Every activation-layer gauge WM evictions, breakthroughs, embedder
// breaker state, context drift, and now the Hebbian counters was therefore
// invisible unless the soul happened to be running and its ISEs were read back
// out of the store. Diagnosing the activation layer required a working soul,
// which is exactly backwards: the lower layer should be observable on its own.
// This review needed it to verify link formation and could not get at it. One
// line of plumbing, and the whole activation layer becomes directly diagnosable.
fn route_act_stats(method: String, path: String, body: String) -> String {
engram_act_stats_json()
}
// route_text_health GET /api/text-health
// (2026-08-08 self-review) The daily census half of the text-integrity gauge.
// Today's review found that the JSON parser had been replacing every \uXXXX
// escape with a literal '?' for at least two months: 3,119 of 4,081
// non-telemetry nodes (76%) were damaged, including the self traversal root
// and every values node, and NOTHING detected it because every gauge in the
// system measured whether the machinery was running, and none measured whether
// the text it carried was intact. No snapshot on disk predates the damage, so
// it cannot be undone; it can only be made impossible to repeat quietly.
//
// The parser is fixed. This route is the standing check: `damaged` should now
// hold flat at its historical floor and never climb. `write_damaged` (also on
// the heartbeat as txt_damaged) is the live regression signal non-zero means
// a write path is mangling text right now.
fn route_text_health(method: String, path: String, body: String) -> String {
engram_text_health_json()
}
// (2026-07-18 self-review) Scoping sweep: `let` inside an if-block creates an
// inner scope only it does NOT mutate the outer binding (documented with
// evidence in awareness.el, 2026-05-25). Every default/reassignment below used
// that broken pattern, so defaults never applied: nodes were created with
// node_type="" and salience=0.0, /api/search and /api/activate ALWAYS ran with
// q="" regardless of input, edges defaulted to relation=""/weight=0.0, and
// save/load with no "path" hit engram_save(""). Rewritten to the
// `let x = if cond { a } else { b }` expression form (the pattern the newer
// routes route_emit_ise/route_capture_knowledge already use correctly).
// persist_canonical save the canonical snapshot after a durable write.
//
// WHY (2026-07-22 self-review): the 2026-07-21 fix correctly stopped READ
// routes from writing the canonical snapshot.json but nothing was left
// that saved it on WRITE. Every mutation (node create, edge create,
// knowledge capture, forget, merge) lived only in RAM until someone POSTed
// /api/save manually; a process restart silently discarded everything since
// the last manual save. Observed live: two engram restarts during the
// 2026-07-22 review reverted the store to a ~17h-old snapshot, destroying
// same-day writes. Reads must never write the canonical; writes must always
// persist it. ISE telemetry is deliberately excluded (48h-pruned, loss-
// tolerant, ~2/min snapshotting the whole store per heartbeat is waste;
// any durable write that follows persists the pruning too).
fn persist_canonical() -> Int {
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
// (2026-08-10 self-review) This returned a hardcoded 1, which made every
// caller's `let saved: Int = persist_canonical()` a dead variable six
// durable write paths each believed they had confirmation of a successful
// canonical persist and none of them had any. Propagate the real result.
return engram_save(dir + "/snapshot.json")
}
// INCOMPLETE-ROUTE FIX (2026-07-24 self-review): this route silently dropped
// label, importance, tier, and tags engram_node() defaults label to content
// and importance to 0.5, so every node created over HTTP lost its metadata.
// Observed live: the soul's boot-counter write-back landed with
// label="soul:boot_count:99" (content), importance 0.5, no tags. Honor the
// full field set via engram_node_full when any of them is supplied.
// PRESENCE-AWARE DEFAULTS (2026-08-01 self-review): the old pattern
// `if x == 0.0 { default }` made a legitimate 0.0 unrepresentable a caller
// setting salience/importance/weight to zero silently got 0.5. json_get_raw
// returns "" when the key is ABSENT and the raw token when present, so
// absence and zero are now distinguishable. Also: confidence was hardcoded
// to 1.0 regardless of input every HTTP-created node claimed full
// epistemic confidence. Now honored from the payload (default 1.0).
fn route_create_node(method: String, path: String, body: String) -> String {
let content: String = json_get_string(body, "content")
let nt_raw: String = json_get_string(body, "node_type")
let node_type: String = if str_eq(nt_raw, "") { "Memory" } else { nt_raw }
let sal_present: String = json_get_raw(body, "salience")
let salience: Float = if str_eq(sal_present, "") { 0.5 } else { json_get_float(body, "salience") }
let label_raw: String = json_get_string(body, "label")
let label: String = if str_eq(label_raw, "") { content } else { label_raw }
let imp_present: String = json_get_raw(body, "importance")
let importance: Float = if str_eq(imp_present, "") { 0.5 } else { json_get_float(body, "importance") }
let conf_present: String = json_get_raw(body, "confidence")
let confidence: Float = if str_eq(conf_present, "") { 1.0 } else { json_get_float(body, "confidence") }
let tier_raw: String = json_get_string(body, "tier")
let tier: String = if str_eq(tier_raw, "") { "Working" } else { tier_raw }
let tags: String = json_get_string(body, "tags")
// NO el_from_float WRAPPER (2026-08-01 self-review): salience/importance/
// confidence are already Float (el_val_t) values json_get_float and
// Float literals both encode. Wrapping them in el_from_float AGAIN
// reinterpreted the boxed bits as a raw double, producing garbage that
// failed engram_decode_score's range check and clamped every HTTP-created
// node to defaults (salience 0.9 in 0.5 stored; confidence 0.6 in → 1.0
// stored verified live). route_emit_ise always passed Floats bare and
// its 0.3/0.3/0.8 stored correctly; this call now does the same.
let id: String = engram_node_full(
content, node_type, label,
salience, importance, confidence,
tier, tags
)
let saved: Int = persist_canonical()
let node_type: String = json_get_string(body, "node_type")
if str_eq(node_type, "") { let node_type = "Memory" }
let salience: Float = json_get_float(body, "salience")
if salience == 0.0 { let salience = 0.5 }
let id: String = engram_node(content, node_type, salience)
"{\"id\":\"" + id + "\",\"content\":\"" + content + "\",\"node_type\":\"" + node_type + "\"}"
}
@@ -202,14 +107,13 @@ fn route_scan_nodes(method: String, path: String, body: String) -> String {
}
// route_scan_edges bulk export of all edges as a JSON array. Implemented
// via engram_save fs_read of a SCRATCH export path. (2026-07-21 self-review:
// previously this saved over the canonical snapshot.json on every GET if the
// process ever booted with a partial/empty store, the first read request
// clobbered the good snapshot. Read routes must never write the canonical path.)
// via engram_save fs_read of the canonical on-disk snapshot, which the
// runtime keeps in lockstep with the in-memory graph. Live against the
// running graph, not a stale export.
fn route_scan_edges(method: String, path: String, body: String) -> String {
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
let snap_path: String = dir + "/.scan-export.json"
let dir: String = env("ENGRAM_DATA_DIR")
if str_eq(dir, "") { let dir = "/tmp/engram" }
let snap_path: String = dir + "/snapshot.json"
engram_save(snap_path)
let snap: String = fs_read(snap_path)
if str_eq(snap, "") { return "[]" }
@@ -222,88 +126,43 @@ fn route_scan_edges(method: String, path: String, body: String) -> String {
}
fn route_search(method: String, path: String, body: String) -> String {
let q: String = if str_eq(method, "GET") { query_param(path, "q") } else { json_get_string(body, "query") }
let lim_url: Int = query_int(path, "limit", 0)
let lim_body: Int = json_get_int(body, "limit")
let lim_either: Int = if lim_url > 0 { lim_url } else { lim_body }
let limit: Int = if lim_either > 0 { lim_either } else { 20 }
let q: String = ""
if str_eq(method, "GET") {
let q = query_param(path, "q")
} else {
let q = json_get_string(body, "query")
}
let limit: Int = query_int(path, "limit", 20)
if limit == 0 { let limit = json_get_int(body, "limit") }
if limit == 0 { let limit = 20 }
return engram_search_json(q, limit)
}
fn route_activate(method: String, path: String, body: String) -> String {
let q: String = if str_eq(method, "GET") { query_param(path, "q") } else { json_get_string(body, "query") }
// Guard: engram_activate with an empty query matches zero seeds, which
// zeroes ALL carried working-memory weights (documented in awareness.el
// perceive()). Never let an empty activation through to wipe WM.
if str_eq(q, "") { return err_json("missing query") }
let d_raw: Int = if str_eq(method, "GET") { query_int(path, "depth", 3) } else { json_get_int(body, "depth") }
let depth: Int = if d_raw > 0 { d_raw } else { 3 }
let q: String = ""
let depth: Int = 3
if str_eq(method, "GET") {
let q = query_param(path, "q")
let depth = query_int(path, "depth", 3)
} else {
let q = json_get_string(body, "query")
let bd: Int = json_get_int(body, "depth")
if bd > 0 { let depth = bd }
}
return "{\"results\":" + engram_activate_json(q, depth) + "}"
}
fn route_create_edge(method: String, path: String, body: String) -> String {
let from_id: String = json_get_string(body, "from_id")
let to_id: String = json_get_string(body, "to_id")
let rel_raw: String = json_get_string(body, "relation")
let relation: String = if str_eq(rel_raw, "") { "associates" } else { rel_raw }
// Presence-aware (2026-08-01): weight 0.0 is a legitimate edge weight
// (dormant association); only default when the key is absent.
let w_present: String = json_get_raw(body, "weight")
let weight: Float = if str_eq(w_present, "") { 0.5 } else { json_get_float(body, "weight") }
let relation: String = json_get_string(body, "relation")
if str_eq(relation, "") { let relation = "associates" }
let weight: Float = json_get_float(body, "weight")
if weight == 0.0 { let weight = 0.5 }
engram_connect(from_id, to_id, weight, relation)
let saved: Int = persist_canonical()
"{\"ok\":true,\"from_id\":\"" + from_id + "\",\"to_id\":\"" + to_id + "\",\"relation\":\"" + relation + "\"}"
}
// route_create_edges_batch POST /api/edges/batch {"edges":[{from_id,to_id,relation,weight}, ...]}
//
// WHY THIS EXISTS (2026-08-07 self-review). persist_canonical() writes the
// FULL canonical snapshot 60MB at current graph size and route_create_edge
// calls it once per edge. That is correct for the interactive one-edge case and
// ruinous for any bulk write: the soul's Hebbian consolidation path delivers
// ~14 associations per 8-minute heartbeat, which through the single-edge route
// would be ~840MB of disk writes per beat, ~150GB/day, to persist 14 edges.
//
// The fix is not to weaken durability it is to make the unit of durability
// the BATCH. Connect every edge, then snapshot exactly once. Same guarantee
// (nothing acknowledged is lost to a restart), 1/N the writes. Empty or
// malformed entries are skipped rather than aborting the batch: a consolidation
// payload is best-effort by design, and one bad id should not cost the other 13.
//
// Returns the accepted count so the caller can tell delivery from silence.
fn route_create_edges_batch(method: String, path: String, body: String) -> String {
let arr: String = json_get_raw(body, "edges")
if str_eq(arr, "") { return err_json("missing edges array") }
let n: Int = json_array_len(arr)
if n == 0 { return "{\"ok\":true,\"accepted\":0,\"skipped\":0}" }
let i: Int = 0
let accepted: Int = 0
let skipped: Int = 0
while i < n {
let item: String = json_array_get(arr, i)
let from_id: String = json_get_string(item, "from_id")
let to_id: String = json_get_string(item, "to_id")
if str_eq(from_id, "") || str_eq(to_id, "") {
let skipped = skipped + 1
} else {
let rel_raw: String = json_get_string(item, "relation")
let relation: String = if str_eq(rel_raw, "") { "associates" } else { rel_raw }
let w_present: String = json_get_raw(item, "weight")
let weight: Float = if str_eq(w_present, "") { 0.5 } else { json_get_float(item, "weight") }
engram_connect(from_id, to_id, weight, relation)
let accepted = accepted + 1
}
let i = i + 1
}
// ONE snapshot for the whole batch the entire point of this route.
// Skip it when nothing was accepted: an all-malformed payload must not
// trigger a 60MB write.
if accepted > 0 {
let saved: Int = persist_canonical()
}
return "{\"ok\":true,\"accepted\":" + int_to_str(accepted) + ",\"skipped\":" + int_to_str(skipped) + "}"
}
fn route_neighbors(method: String, path: String, body: String) -> String {
let id: String = extract_id(path, "/api/neighbors/")
if str_eq(id, "") { return err_json("missing id") }
@@ -315,7 +174,6 @@ fn route_strengthen(method: String, path: String, body: String) -> String {
let id: String = json_get_string(body, "node_id")
if str_eq(id, "") { return err_json("missing node_id") }
engram_strengthen(id)
let saved: Int = persist_canonical()
ok_json()
}
@@ -323,84 +181,33 @@ fn route_forget(method: String, path: String, body: String) -> String {
let id: String = extract_id(path, "/api/nodes/")
if str_eq(id, "") { return err_json("missing id") }
engram_forget(id)
let saved: Int = persist_canonical()
ok_json()
}
fn route_save(method: String, path: String, body: String) -> String {
let p_raw: String = json_get_string(body, "path")
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
let p: String = if str_eq(p_raw, "") { dir + "/snapshot.json" } else { p_raw }
// (2026-08-10 self-review) engram_save returns 0 on an empty path and the
// route discarded it, so the response was a literal "ok":true regardless
// of whether anything was written. Report the actual result AND the counts
// that were supposed to have been written the same move that made
// route_health honest on 2026-08-01. A caller can now tell "saved 13k
// nodes" from "saved nothing and said ok".
let sv: Int = engram_save(p)
let sv_ok: String = if sv == 0 { "false" } else { "true" }
"{\"ok\":" + sv_ok + ",\"path\":\"" + p + "\",\"node_count\":" + int_to_str(engram_node_count()) + ",\"edge_count\":" + int_to_str(engram_edge_count()) + "}"
let p: String = json_get_string(body, "path")
if str_eq(p, "") {
let dir: String = env("ENGRAM_DATA_DIR")
if str_eq(dir, "") { let dir = "/tmp/engram" }
let p = dir + "/snapshot.json"
}
engram_save(p)
"{\"ok\":true,\"path\":\"" + p + "\"}"
}
fn route_load(method: String, path: String, body: String) -> String {
let p_raw: String = json_get_string(body, "path")
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
let p: String = if str_eq(p_raw, "") { dir + "/snapshot.json" } else { p_raw }
// (2026-08-10 self-review) This was a stub response over the single most
// destructive operation in the server. engram_load returns 0 on an empty
// path, an unopenable file, a zero-length file, or malloc failure and
// this route answered ok_json() in every one of those cases.
//
// Precise failure shape (el_runtime.c:9890): the fopen guard runs BEFORE
// the store reset, so a MISSING path is genuinely safe it returns 0 with
// the graph intact. The dangerous case is a readable-but-malformed file:
// the reset loop frees every node and edge FIRST, then parses, so a
// truncated or non-snapshot JSON leaves a hollow store and the caller
// was told "ok":true. With 37 GB of stale dated snapshots sitting in the
// data dir as tempting restore targets, "restore reported success and
// silently emptied the graph" is a live risk, not a hypothetical one.
//
// Fix: surface the return value AND the resulting counts. node_count=0
// after a load is the unambiguous hollow-store signal (same convention
// route_health adopted 2026-08-01). Callers can now verify a restore
// instead of trusting it.
let ld: Int = engram_load(p)
let ld_ok: String = if ld == 0 { "false" } else { "true" }
let nc_after: Int = engram_node_count()
let hollow: String = if nc_after == 0 { "true" } else { "false" }
"{\"ok\":" + ld_ok + ",\"path\":\"" + p + "\",\"node_count\":" + int_to_str(nc_after) + ",\"edge_count\":" + int_to_str(engram_edge_count()) + ",\"hollow\":" + hollow + "}"
}
// (2026-08-01 self-review) Health previously returned a hardcoded literal
// it reported "ok" even when the snapshot failed to load and the store was
// empty. Now reports live counts so a monitor can distinguish "up and
// loaded" from "up and hollow" (node_count=0 after boot = failed load).
fn route_health(method: String, path: String, body: String) -> String {
"{\"status\":\"ok\",\"engine\":\"engram-runtime-native\",\"node_count\":" + int_to_str(engram_node_count()) + ",\"edge_count\":" + int_to_str(engram_edge_count()) + "}"
}
// route_embed_backfill GET/POST /api/embed-backfill?n=48
//
// (2026-07-25 self-review) The lazy embedding backfill runs only inside
// engram_activate, and nothing in production calls /api/activate on this
// store the soul's curiosity loop activates its own in-process graph.
// After a restart from a snapshot without vectors, embedded_count stalled
// at 93/12175 and would never recover. This route lets the soul's
// heartbeat pump the backfill explicitly (48/min clears a 12k backlog in
// ~4h). Persists the canonical snapshot whenever new vectors were
// generated the 2026-07-25 regression happened precisely because 3747
// in-RAM embeddings were never snapshotted before a restart. Self-
// limiting: once coverage is full, embedded=0 and no save occurs.
fn route_embed_backfill(method: String, path: String, body: String) -> String {
let n: Int = query_int(path, "n", 32)
let result: String = engram_embed_backfill(n)
let done: Float = json_get_float(result, "embedded")
if done > 0.0 {
let saved: Int = persist_canonical()
let p: String = json_get_string(body, "path")
if str_eq(p, "") {
let dir: String = env("ENGRAM_DATA_DIR")
if str_eq(dir, "") { let dir = "/tmp/engram" }
let p = dir + "/snapshot.json"
}
return result
engram_load(p)
ok_json()
}
fn route_health(method: String, path: String, body: String) -> String {
"{\"status\":\"ok\",\"engine\":\"engram-runtime-native\"}"
}
// route_sync return a snapshot of non-ISE/non-Working nodes for the soul daemon
@@ -416,45 +223,15 @@ fn route_embed_backfill(method: String, path: String, body: String) -> String {
// (it skips nodes already present by ID). Auth-exempt: same-host internal call.
// (2026-06-27 self-review: added this route to fix silent 10-min sync failures)
fn route_sync(method: String, path: String, body: String) -> String {
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
// 2026-07-21 self-review: export to a scratch path, never the canonical
// snapshot.json read routes must not be able to clobber the good snapshot.
let snap_path: String = dir + "/.sync-export.json"
let dir: String = env("ENGRAM_DATA_DIR")
if str_eq(dir, "") { let dir = "/tmp/engram" }
let snap_path: String = dir + "/snapshot.json"
engram_save(snap_path)
let snap: String = fs_read(snap_path)
// 2026-08-02 self-review: this used to return {"nodes":[],"edges":[]} when
// the export/read failed. The soul's sync_ok test (awareness.el) only
// checks for "" and "{}", so that placeholder PASSED as a healthy sync:
// soul.last_sync_ok_ts got stamped, sync_age_ms stayed green, the
// sync_empty warn ISE never fired, and engram_sync reported added:0
// forever. A totally broken sync was indistinguishable from a quiet
// healthy one the exact failure class this route was added to fix in
// the first place (see 2026-06-27 note above). Return a real error so the
// failure is loud on both sides.
if str_eq(snap, "") { return err_json("sync export failed: snapshot unreadable") }
if str_eq(snap, "") { return "{\"nodes\":[],\"edges\":[]}" }
return snap
}
// route_load_merge POST /api/load-merge {"path": "..."} merge a snapshot
// file into the live store WITHOUT resetting it (engram_load_merge skips nodes
// already present by id). Added 2026-07-21 self-review to restore the 244 kn-
// identity Knowledge nodes lost from the snapshot lineage between 05-13 and
// 07-13. Requires an explicit path: refuses to run without one so it can never
// be triggered accidentally against a default.
fn route_load_merge(method: String, path: String, body: String) -> String {
let p: String = json_get_string(body, "path")
if str_eq(p, "") { return err_json("path is required") }
if str_eq(fs_read(p), "") { return err_json("file missing or empty") }
let before_n: Int = engram_node_count()
let before_e: Int = engram_edge_count()
engram_load_merge(p)
let added_n: Int = engram_node_count() - before_n
let added_e: Int = engram_edge_count() - before_e
let saved: Int = persist_canonical()
"{\"ok\":true,\"nodes_added\":" + int_to_str(added_n) + ",\"edges_added\":" + int_to_str(added_e) + ",\"node_count\":" + int_to_str(engram_node_count()) + "}"
}
// route_emit_ise write an InternalStateEvent node from the soul daemon.
//
// Endpoint: POST /api/neuron/state-events
@@ -468,20 +245,10 @@ fn route_load_merge(method: String, path: String, body: String) -> String {
//
// Salience/importance set to match engram_node_full ISE defaults used by the
// in-process fallback path in awareness.el (salience=0.3, importance=0.3,
// confidence=0.8, tier=Episodic).
// confidence=0.8, tier=Episodic). High temporal_decay_rate (1.617) ISEs
// are inherently transient; they should decay faster than structural knowledge.
// (2026-06-26 self-review: added this route after discovering ise_post was
// silently failing the soul posts here but the endpoint didn't exist.)
//
// Retention (2026-07-16 self-review): an earlier comment here claimed ISEs
// got temporal_decay_rate=1.617 that was never implemented (engram_node_full
// hardcodes 0.0), and per-node decay only dampens activation anyway; it never
// removes nodes. By 2026-07-16 ISEs were 75% of the store (10,175 of 13,522
// nodes, ~4,300/day, unbounded). ISEs are already WM-excluded in
// engram_activate, so the fix is retention, not decay: every insert calls
// engram_prune_telemetry(), a single O(nodes+edges) compaction pass that
// removes ISEs older than ENGRAM_ISE_RETENTION_MS (default 48h), protecting
// "session-start" labels and self_review events as durable history. At
// ~3 ISEs/min this bounds telemetry at ~8.6k nodes instead of growing forever.
fn route_emit_ise(method: String, path: String, body: String) -> String {
let content: String = json_get_string(body, "content")
if str_eq(content, "") { return err_json("missing content") }
@@ -493,86 +260,9 @@ fn route_emit_ise(method: String, path: String, body: String) -> String {
sal, imp, conf,
"Episodic", "[\"internal-state\",\"InternalStateEvent\"]"
)
let ret_raw: String = env("ENGRAM_ISE_RETENTION_MS")
let ret_ms: Int = if str_eq(ret_raw, "") { 172800000 } else { str_to_int(ret_raw) }
let pruned: Int = engram_prune_telemetry(ret_ms)
"{\"ok\":true,\"id\":\"" + id + "\",\"pruned\":" + int_to_str(pruned) + "}"
}
// Knowledge capture
//
// route_capture_knowledge direct Knowledge-node capture over HTTP.
//
// Endpoint: POST /api/neuron/knowledge/capture (auth required: "_auth" in body)
// Body: {"content": "...", "title": "...", "category": "...",
// "tier": "note|lesson|canonical", "tags": [...], "project": "...",
// "_auth": "<key>"}
//
// WHY (2026-07-15 self-review): the world-ingestor integrator was designed
// against this endpoint (its MCP-unavailable fallback), but the route never
// existed every direct push 404'd, and because the auth gate ran before
// routing, the failure surfaced as {"error":"unauthorized"} and was
// misdiagnosed for two weeks while world knowledge silently dropped.
// POST /api/nodes was no substitute: it discards label/tags/tier, which
// makes captured knowledge invisible to tag-scoped search and curiosity.
//
// The incoming knowledge tier (note/lesson/canonical) is preserved as a
// "tier:<x>" tag rather than mapped onto Engram's cognitive tiers Knowledge
// nodes land in Semantic (stable reference), and the epistemic tier stays
// queryable without inventing a lossy mapping.
fn route_capture_knowledge(method: String, path: String, body: String) -> String {
let content: String = json_get_string(body, "content")
if str_eq(content, "") { return err_json("missing content") }
let title: String = json_get_string(body, "title")
let label: String = if str_eq(title, "") { str_slice(content, 0, 60) } else { title }
let category_raw: String = json_get_string(body, "category")
let category: String = if str_eq(category_raw, "") { "other" } else { category_raw }
let ktier_raw: String = json_get_string(body, "tier")
let ktier: String = if str_eq(ktier_raw, "") { "note" } else { ktier_raw }
let project: String = json_get_string(body, "project")
let tags_raw: String = json_get_raw(body, "tags")
let tags_base: String = if str_eq(tags_raw, "") { "[]" } else { tags_raw }
// Merge category/tier/project markers into the tag array. Search matches
// against the tags string, so these make captures findable by facet.
let base_len: Int = str_len(tags_base)
let head: String = str_slice(tags_base, 0, base_len - 1)
let sep: String = if str_eq(head, "[") { "" } else { "," }
let safe_cat: String = str_replace(category, "\"", "'")
let safe_tier: String = str_replace(ktier, "\"", "'")
let safe_proj: String = str_replace(project, "\"", "'")
let proj_tag: String = if str_eq(safe_proj, "") { "" } else { ",\"project:" + safe_proj + "\"" }
let tags: String = head + sep + "\"category:" + safe_cat + "\",\"tier:" + safe_tier + "\"" + proj_tag + "]"
let sal: Float = 0.5
let imp: Float = 0.5
let conf: Float = 0.9
let id: String = engram_node_full(
content, "Knowledge", label,
sal, imp, conf,
"Semantic", tags
)
let saved: Int = persist_canonical()
"{\"ok\":true,\"id\":\"" + id + "\"}"
}
// route_similarity GET /api/similarity?a=<id>&b=<id>
//
// (2026-08-01 self-review) engram_cosine_sim was added 2026-07-24
// (bl-b2d1c944) with the stated purpose of exposing semantic distance to
// "EL code and the introspection API" but it had ZERO callers anywhere:
// no route, no soul-daemon use. The activation path uses embeddings
// internally (semantic seeding, Pass-2 additive term), but there was no way
// to probe pairwise node similarity from outside. This closes that: cosine
// in [-1,1], or -2 when either node is missing or not yet embedded (so
// "not comparable" is distinguishable from "genuinely orthogonal" 0.0).
fn route_similarity(method: String, path: String, body: String) -> String {
let a: String = query_param(path, "a")
let b: String = query_param(path, "b")
if str_eq(a, "") { return err_json("missing a") }
if str_eq(b, "") { return err_json("missing b") }
let sim: Float = engram_cosine_sim(a, b)
"{\"a\":\"" + a + "\",\"b\":\"" + b + "\",\"cosine\":" + float_to_str(sim) + "}"
}
// Auth
fn check_auth_ok(method: String, body: String) -> Bool {
@@ -609,22 +299,10 @@ fn handle_request(method: String, path: String, body: String) -> String {
return err_json("unauthorized")
}
// Knowledge capture (auth enforced above; the world-ingestor integrator
// and any headless session without MCP push knowledge through this)
if str_eq(method, "POST") && str_eq(clean, "/api/neuron/knowledge/capture") {
return route_capture_knowledge(method, path, body)
}
// Stats
if str_eq(method, "GET") && (str_eq(clean, "/api/stats") || str_eq(clean, "/stats")) {
return route_stats(method, path, body)
}
if str_eq(method, "GET") && (str_eq(clean, "/api/act-stats") || str_eq(clean, "/act-stats")) {
return route_act_stats(method, path, body)
}
if str_eq(method, "GET") && (str_eq(clean, "/api/text-health") || str_eq(clean, "/text-health")) {
return route_text_health(method, path, body)
}
// Nodes
if str_eq(method, "POST") && (str_eq(clean, "/api/nodes") || str_eq(clean, "/nodes")) {
@@ -647,13 +325,6 @@ fn handle_request(method: String, path: String, body: String) -> String {
if str_eq(method, "POST") && (str_eq(clean, "/api/edges") || str_eq(clean, "/edges")) {
return route_create_edge(method, path, body)
}
// Batch edge write one snapshot for the whole payload. Must be tested
// BEFORE nothing else claims it; the exact-match on "/api/edges" above
// does not catch "/api/edges/batch", so order is not load-bearing here,
// but keeping the two adjacent keeps them from drifting apart.
if str_eq(method, "POST") && (str_eq(clean, "/api/edges/batch") || str_eq(clean, "/edges/batch")) {
return route_create_edges_batch(method, path, body)
}
if str_eq(method, "GET") && str_starts_with(clean, "/api/neighbors/") {
return route_neighbors(method, path, body)
}
@@ -684,55 +355,27 @@ fn handle_request(method: String, path: String, body: String) -> String {
if str_eq(method, "POST") && (str_eq(clean, "/api/load") || str_eq(clean, "/load")) {
return route_load(method, path, body)
}
if str_eq(method, "POST") && (str_eq(clean, "/api/load-merge") || str_eq(clean, "/load-merge")) {
return route_load_merge(method, path, body)
}
// Sync soul daemon periodic pull of non-ISE knowledge into in-process graph
if str_eq(method, "GET") && str_eq(clean, "/api/sync") {
return route_sync(method, path, body)
}
// Embedding backfill pumped by the soul heartbeat (2026-07-25)
if str_eq(clean, "/api/embed-backfill") {
return route_embed_backfill(method, path, body)
}
// Semantic similarity probe (2026-08-01)
if str_eq(method, "GET") && str_starts_with(clean, "/api/similarity") {
return route_similarity(method, path, body)
}
"{\"error\":\"not found\",\"path\":\"" + clean + "\"}"
}
// Entry
let bind_raw: String = env("ENGRAM_BIND")
let bind_str: String = if str_eq(bind_raw, "") { ":8742" } else { bind_raw }
let bind_str: String = env("ENGRAM_BIND")
if str_eq(bind_str, "") { let bind_str = ":8742" }
let port: Int = parse_port(bind_str)
// On startup, try to load any existing snapshot (best effort).
let data_dir_raw: String = env("ENGRAM_DATA_DIR")
let data_dir: String = if str_eq(data_dir_raw, "") { "/tmp/engram" } else { data_dir_raw }
let data_dir: String = env("ENGRAM_DATA_DIR")
if str_eq(data_dir, "") { let data_dir = "/tmp/engram" }
let snapshot_path: String = data_dir + "/snapshot.json"
engram_load(snapshot_path)
// 2026-07-21 self-review boot guard: if the snapshot file has content but the
// load produced 0 nodes, something is wrong (corrupt file / parse failure).
// Preserve the evidence and warn loudly and since read routes no longer write
// the canonical path, a bad boot can no longer clobber the good snapshot.
let boot_snap: String = fs_read(snapshot_path)
if !str_eq(boot_snap, "") {
if engram_node_count() == 0 {
println("[engram] WARNING: snapshot.json is non-empty but load produced 0 nodes — preserving copy at snapshot.failed-load.json")
fs_write(data_dir + "/snapshot.failed-load.json", boot_snap)
} else {
// Good load: keep a boot-time backup of the snapshot as loaded.
fs_write(data_dir + "/snapshot.boot-backup.json", boot_snap)
}
}
println("[engram] runtime-native graph engine")
println("[engram] data_dir=" + data_dir)
println("[engram] node_count=" + int_to_str(engram_node_count()))
+187 -413
View File
@@ -82,8 +82,14 @@ static _Thread_local ElArena _tl_arena = {NULL, 0, 0};
static _Thread_local int _tl_arena_active = 0;
/* Binary-safe fs_read length — set by fs_read, consumed by http_send_response.
* Allows serving PNGs and other binary files without strlen truncation. */
static _Thread_local size_t _tl_fs_read_len = 0;
* Allows serving PNGs and other binary files without strlen truncation.
* PAIRED with the buffer pointer it describes: the length may only be applied
* to the exact buffer fs_read returned. Without the pairing, any handler that
* fs_read a file and then WRAPPED it into a larger response had that response
* truncated to the file's length (Content-Length lied AND the send stopped
* short) the safety-contact onboarding trap, 2026-07-17. */
static _Thread_local size_t _tl_fs_read_len = 0;
static _Thread_local const char* _tl_fs_read_buf = NULL;
static void el_arena_track(char* p) {
if (!_tl_arena_active || !p) return;
@@ -101,6 +107,8 @@ static void el_arena_track(char* p) {
void el_request_start(void) {
_tl_arena.count = 0;
_tl_arena_active = 1;
_tl_fs_read_len = 0; /* never let a previous request's file length */
_tl_fs_read_buf = NULL; /* leak into this response's byte accounting */
}
/* Called by http_worker after the El handler returns and the response is sent.
@@ -1484,11 +1492,14 @@ static void http_send_response(int fd, const char* body) {
}
const char* eff_body = is_envelope ? env_body : body;
/* Use the real byte count from fs_read if available (handles binary files
* with embedded null bytes PNG, WOFF2, etc.). Fall back to strlen for
* normal text/JSON responses where _tl_fs_read_len is 0. */
size_t blen = (_tl_fs_read_len > 0) ? _tl_fs_read_len : strlen(eff_body);
/* Use the real byte count from fs_read ONLY when this body IS the exact
* buffer fs_read returned (binary files with embedded null bytes PNG,
* WOFF2, etc.). Any other body wrapped, enveloped, or derived must be
* measured with strlen, or it is truncated/over-read to the file's size. */
size_t blen = (_tl_fs_read_len > 0 && eff_body == _tl_fs_read_buf)
? _tl_fs_read_len : strlen(eff_body);
_tl_fs_read_len = 0; /* consume — one-shot per response */
_tl_fs_read_buf = NULL;
int head_only = _tl_http_head_only;
JsonBuf hdrs; jb_init(&hdrs);
@@ -1545,17 +1556,6 @@ typedef struct {
#endif
} HttpWorkerArg;
/* Forward declarations for the loopback/API-key hardening helpers defined
* further down. Without these, http_worker's calls below were implicit
* declarations and the later `static` definitions conflicted with them this
* file did not compile at all. (2026-08-08 self-review: the hardening work
* they belong to had been sitting uncommitted in the working tree since
* 2026-07-15 in exactly this non-building state, which is presumably why it
* was never committed. Adding the two prototypes is the whole fix.) */
static int el_http_request_authorized(const char* method, const char* path,
const char* hdr_block);
static void el_http_send_401(int fd);
static void* http_worker(void* arg) {
HttpWorkerArg* a = (HttpWorkerArg*)arg;
#ifdef _WIN32
@@ -1564,13 +1564,8 @@ static void* http_worker(void* arg) {
int fd = a->fd;
#endif
free(a);
char *method = NULL, *path = NULL, *body = NULL, *hdr_block = NULL;
if (http_read_request(fd, &method, &path, &body, &hdr_block) == 0
&& !el_http_request_authorized(method, path, hdr_block)) {
/* Loopback hardening: EL_HTTP_AUTH_KEY is set and this request lacks the
* matching X-Neuron-Auth header refuse before it reaches any handler. */
el_http_send_401(fd);
} else if (method != NULL) {
char *method = NULL, *path = NULL, *body = NULL;
if (http_read_request(fd, &method, &path, &body, NULL) == 0) {
http_handler_fn h = http_lookup_active();
char* response = NULL;
/* HEAD: dispatch as GET so existing handlers respond with the same
@@ -1584,11 +1579,22 @@ static void* http_worker(void* arg) {
const char* rs = EL_CSTR(r);
/* Copy response out BEFORE arena teardown.
* For binary files, _tl_fs_read_len holds the real byte count
* use memcpy instead of strdup so null bytes are preserved. */
size_t rlen = _tl_fs_read_len > 0 ? _tl_fs_read_len : (rs ? strlen(rs) : 0);
* use memcpy instead of strdup so null bytes are preserved.
* The stored length applies ONLY when the response IS the exact
* fs_read buffer; a wrapped/derived response must use strlen or
* it gets truncated (or over-read) to the file's length. */
size_t rlen;
if (_tl_fs_read_len > 0 && rs && rs == _tl_fs_read_buf) {
rlen = _tl_fs_read_len; /* raw file bytes — binary-safe */
} else {
rlen = rs ? strlen(rs) : 0;
_tl_fs_read_len = 0; /* hint doesn't describe this body */
_tl_fs_read_buf = NULL;
}
response = malloc(rlen + 1);
if (response && rs) { memcpy(response, rs, rlen); response[rlen] = '\0'; }
else if (response) { response[0] = '\0'; }
if (_tl_fs_read_len > 0) _tl_fs_read_buf = response; /* hint follows the copy */
} else {
response = el_strdup_persist("el-runtime: no http handler registered");
}
@@ -1598,7 +1604,7 @@ static void* http_worker(void* arg) {
_tl_http_head_only = 0;
free(response);
}
free(method); free(path); free(body); free(hdr_block);
free(method); free(path); free(body);
el_closesocket(fd);
/* release a slot */
pthread_mutex_lock(&_http_conn_mu);
@@ -1608,108 +1614,6 @@ static void* http_worker(void* arg) {
return NULL;
}
/* ── loopback lock + local API-key auth (shipped desktop hardening) ────────
* Both controls are OFF by default (their env vars unset), so dev, self-host,
* and server builds behave exactly as before. The shipped macOS launcher
* neuron-daemons.sh sets them so a customer's soul is neither reachable from
* other machines on the LAN nor callable by other local users/processes
* without the per-install key held in the login Keychain:
*
* EL_HTTP_BIND_HOST=127.0.0.1 -> bind loopback only (el_http_apply_bind_addr)
* EL_HTTP_AUTH_KEY=<per-install> -> require "X-Neuron-Auth: <key>" per request
*/
/* Set the listen address on the dual-stack (AF_INET6, V6ONLY=0) socket. Default
* is in6addr_any (all interfaces) unchanged. When EL_HTTP_BIND_HOST names a
* loopback ("127.0.0.1", "localhost", "loopback", or "::1") we bind the IPv4-
* mapped IPv6 loopback ::ffff:127.0.0.1: on a V6ONLY=0 socket this accepts IPv4
* 127.0.0.1 clients (the desktop app connects there) while refusing every
* off-machine address. */
static void el_http_apply_bind_addr(struct sockaddr_in6* addr) {
const char* h = getenv("EL_HTTP_BIND_HOST");
int loopback = h && *h && (strcmp(h, "127.0.0.1") == 0
|| strcmp(h, "localhost") == 0
|| strcmp(h, "loopback") == 0
|| strcmp(h, "::1") == 0);
if (loopback) {
memset(&addr->sin6_addr, 0, sizeof(addr->sin6_addr));
addr->sin6_addr.s6_addr[10] = 0xff; /* ::ffff:127.0.0.1 */
addr->sin6_addr.s6_addr[11] = 0xff;
addr->sin6_addr.s6_addr[12] = 127;
addr->sin6_addr.s6_addr[15] = 1;
} else {
addr->sin6_addr = in6addr_any;
}
}
/* Human-readable description of the active bind host, for the listen log line. */
static const char* el_http_bind_desc(void) {
const char* h = getenv("EL_HTTP_BIND_HOST");
if (h && *h && (strcmp(h, "127.0.0.1") == 0 || strcmp(h, "localhost") == 0
|| strcmp(h, "loopback") == 0 || strcmp(h, "::1") == 0)) {
return "127.0.0.1 (loopback)";
}
return "[::] (dual-stack)";
}
/* Case-insensitive compare of the first n bytes of a and b. */
static int el_ci_eq_n(const char* a, const char* b, size_t n) {
for (size_t i = 0; i < n; i++) {
unsigned char ca = (unsigned char)a[i], cb = (unsigned char)b[i];
if (tolower(ca) != tolower(cb)) return 0;
}
return 1;
}
/* Return 1 iff the raw header block carries a header named `name` (case-
* insensitive) whose trimmed value equals `want` exactly. */
static int el_http_header_equals(const char* hdr_block, const char* name,
const char* want) {
if (!hdr_block || !name || !want) return 0;
size_t nlen = strlen(name), wlen = strlen(want);
const char* p = hdr_block;
while (*p) {
const char* line_end = strstr(p, "\r\n");
const char* end = line_end ? line_end : p + strlen(p);
const char* colon = memchr(p, ':', (size_t)(end - p));
if (colon && (size_t)(colon - p) == nlen && el_ci_eq_n(p, name, nlen)) {
const char* v = colon + 1;
while (v < end && (*v == ' ' || *v == '\t')) v++;
size_t vlen = (size_t)(end - v);
while (vlen > 0 && (v[vlen - 1] == ' ' || v[vlen - 1] == '\t')) vlen--;
if (vlen == wlen && memcmp(v, want, wlen) == 0) return 1;
}
if (!line_end) break;
p = line_end + 2;
}
return 0;
}
/* Authorize an inbound request. Enforcement is active only when EL_HTTP_AUTH_KEY
* is set; otherwise every request is allowed (dev default). GET/HEAD /health*
* are always allowed so launch-agent liveness probes work without the key. */
static int el_http_request_authorized(const char* method, const char* path,
const char* hdr_block) {
const char* key = getenv("EL_HTTP_AUTH_KEY");
if (!key || !*key) return 1;
if (method && (strcmp(method, "GET") == 0 || strcmp(method, "HEAD") == 0)
&& path && strncmp(path, "/health", 7) == 0) return 1;
return el_http_header_equals(hdr_block, "x-neuron-auth", key);
}
/* Minimal 401 for unauthorized requests — never reaches an EL handler. */
static void el_http_send_401(int fd) {
static const char* body = "{\"error\":\"unauthorized\",\"code\":\"auth_required\"}";
char resp[256];
int n = snprintf(resp, sizeof(resp),
"HTTP/1.1 401 Unauthorized\r\n"
"Content-Type: application/json\r\n"
"Content-Length: %zu\r\n"
"Connection: close\r\n\r\n%s",
strlen(body), body);
if (n > 0) http_send_all(fd, resp, (size_t)n);
}
el_val_t http_serve(el_val_t port, el_val_t handler) {
/* If `handler` looks like a string name, register it as the active handler. */
const char* hname = EL_CSTR(handler);
@@ -1728,13 +1632,13 @@ el_val_t http_serve(el_val_t port, el_val_t handler) {
struct sockaddr_in6 addr;
memset(&addr, 0, sizeof(addr));
addr.sin6_family = AF_INET6;
el_http_apply_bind_addr(&addr);
addr.sin6_addr = in6addr_any;
addr.sin6_port = htons((uint16_t)p);
if (bind(sock, (struct sockaddr*)&addr, sizeof(addr)) < 0) {
perror("bind"); el_closesocket(sock); return 0;
}
if (listen(sock, 64) < 0) { perror("listen"); el_closesocket(sock); return 0; }
fprintf(stderr, "[http] listening on %s port %d\n", el_http_bind_desc(), p);
fprintf(stderr, "[http] listening on [::]:%d (dual-stack)\n", p);
while (1) {
struct sockaddr_in6 cli;
socklen_t clen = sizeof(cli);
@@ -1940,10 +1844,20 @@ static void* http_worker_v2(void* arg) {
el_val_t hmap = http_build_headers_map(hdr_block ? hdr_block : "");
el_val_t r = h(EL_STR(dispatch_method), EL_STR(path), hmap, EL_STR(body));
const char* rs = EL_CSTR(r);
size_t rlen = _tl_fs_read_len > 0 ? _tl_fs_read_len : (rs ? strlen(rs) : 0);
/* Same pairing rule as the v1 worker: the fs_read length is only
* trustworthy for the exact buffer fs_read returned. */
size_t rlen;
if (_tl_fs_read_len > 0 && rs && rs == _tl_fs_read_buf) {
rlen = _tl_fs_read_len; /* raw file bytes — binary-safe */
} else {
rlen = rs ? strlen(rs) : 0;
_tl_fs_read_len = 0; /* hint doesn't describe this body */
_tl_fs_read_buf = NULL;
}
response = malloc(rlen + 1);
if (response && rs) { memcpy(response, rs, rlen); response[rlen] = '\0'; }
else if (response) { response[0] = '\0'; }
if (_tl_fs_read_len > 0) _tl_fs_read_buf = response; /* hint follows the copy */
el_release(hmap);
} else {
response = el_strdup_persist(
@@ -1984,13 +1898,13 @@ el_val_t http_serve_v2(el_val_t port, el_val_t handler) {
struct sockaddr_in6 addr;
memset(&addr, 0, sizeof(addr));
addr.sin6_family = AF_INET6;
el_http_apply_bind_addr(&addr);
addr.sin6_addr = in6addr_any;
addr.sin6_port = htons((uint16_t)p);
if (bind(sock, (struct sockaddr*)&addr, sizeof(addr)) < 0) {
perror("bind"); el_closesocket(sock); return 0;
}
if (listen(sock, 64) < 0) { perror("listen"); el_closesocket(sock); return 0; }
fprintf(stderr, "[http v2] listening on %s port %d\n", el_http_bind_desc(), p);
fprintf(stderr, "[http v2] listening on [::]:%d (dual-stack)\n", p);
while (1) {
struct sockaddr_in6 cli;
socklen_t clen = sizeof(cli);
@@ -2086,13 +2000,13 @@ void http_serve_async(el_val_t port, el_val_t handler) {
struct sockaddr_in6 addr;
memset(&addr, 0, sizeof(addr));
addr.sin6_family = AF_INET6;
el_http_apply_bind_addr(&addr);
addr.sin6_addr = in6addr_any;
addr.sin6_port = htons((uint16_t)p);
if (bind(sock, (struct sockaddr*)&addr, sizeof(addr)) < 0) {
perror("bind"); close(sock); return;
}
if (listen(sock, 64) < 0) { perror("listen"); close(sock); return; }
fprintf(stderr, "[http] async listening on %s port %d\n", el_http_bind_desc(), p);
fprintf(stderr, "[http] async listening on [::]:%d (dual-stack)\n", p);
HttpServeAsyncArg* a = malloc(sizeof(HttpServeAsyncArg));
if (!a) { close(sock); return; }
a->sock = sock;
@@ -2141,6 +2055,7 @@ el_val_t http_response(el_val_t status, el_val_t headers_json, el_val_t body) {
el_val_t fs_read(el_val_t pathv) {
const char* path = EL_CSTR(pathv);
_tl_fs_read_len = 0;
_tl_fs_read_buf = NULL;
if (!path) return el_wrap_str(el_strdup(""));
FILE* f = fopen(path, "rb");
if (!f) return el_wrap_str(el_strdup(""));
@@ -2152,6 +2067,7 @@ el_val_t fs_read(el_val_t pathv) {
size_t got = fread(buf, 1, (size_t)sz, f);
buf[got] = '\0';
_tl_fs_read_len = got; /* store real byte count for binary-safe send */
_tl_fs_read_buf = buf; /* ...valid ONLY for this exact buffer */
fclose(f);
return el_wrap_str(buf);
}
@@ -3257,72 +3173,10 @@ static char* jp_parse_string_raw(JsonParser* jp) {
case 'r': c = '\r'; break;
case 't': c = '\t'; break;
case 'u': {
/* Decode \uXXXX (with surrogate pairs) to UTF-8.
* Ported from lang/releases/v1.0.0-20260501 (2026-08-08
* self-review). This copy carried the identical defect:
* the escape was skipped and a literal '?' emitted, which
* silently destroyed every non-ASCII character in any JSON
* string entering the runtime. Two copies of one parser
* bug is exactly how this class of fault survives, so the
* fix lands in both. See the release copy for the full
* measurement and rationale. */
unsigned cp = 0;
int ok = 1;
for (int i = 0; i < 4; i++) {
if (jp->p >= jp->end) { ok = 0; break; }
char h = *jp->p++;
unsigned d;
if (h >= '0' && h <= '9') d = (unsigned)(h - '0');
else if (h >= 'a' && h <= 'f') d = (unsigned)(h - 'a' + 10);
else if (h >= 'A' && h <= 'F') d = (unsigned)(h - 'A' + 10);
else { ok = 0; break; }
cp = (cp << 4) | d;
}
if (!ok) { c = '?'; break; }
if (cp >= 0xD800 && cp <= 0xDBFF &&
(size_t)(jp->end - jp->p) >= 6 &&
jp->p[0] == '\\' && jp->p[1] == 'u') {
const char* save = jp->p;
unsigned lo = 0; int ok2 = 1;
jp->p += 2;
for (int i = 0; i < 4; i++) {
char h = *jp->p++;
unsigned d;
if (h >= '0' && h <= '9') d = (unsigned)(h - '0');
else if (h >= 'a' && h <= 'f') d = (unsigned)(h - 'a' + 10);
else if (h >= 'A' && h <= 'F') d = (unsigned)(h - 'A' + 10);
else { ok2 = 0; break; }
lo = (lo << 4) | d;
}
if (ok2 && lo >= 0xDC00 && lo <= 0xDFFF)
cp = 0x10000u + ((cp - 0xD800u) << 10) + (lo - 0xDC00u);
else jp->p = save;
}
if (cp >= 0xD800 && cp <= 0xDFFF) cp = 0xFFFD;
char ub[4]; int un;
if (cp < 0x80) {
ub[0] = (char)cp; un = 1;
} else if (cp < 0x800) {
ub[0] = (char)(0xC0 | (cp >> 6));
ub[1] = (char)(0x80 | (cp & 0x3F)); un = 2;
} else if (cp < 0x10000) {
ub[0] = (char)(0xE0 | (cp >> 12));
ub[1] = (char)(0x80 | ((cp >> 6) & 0x3F));
ub[2] = (char)(0x80 | (cp & 0x3F)); un = 3;
} else {
ub[0] = (char)(0xF0 | (cp >> 18));
ub[1] = (char)(0x80 | ((cp >> 12) & 0x3F));
ub[2] = (char)(0x80 | ((cp >> 6) & 0x3F));
ub[3] = (char)(0x80 | (cp & 0x3F)); un = 4;
}
while (len + (size_t)un >= cap) {
cap *= 2;
out = realloc(out, cap);
if (!out) { fputs("el_runtime: out of memory\n", stderr); exit(1); }
}
for (int i = 0; i < un; i++) out[len++] = ub[i];
continue; /* bytes already appended */
/* Skip 4 hex digits; emit '?' as a placeholder */
for (int i = 0; i < 4 && jp->p < jp->end; i++) jp->p++;
c = '?';
break;
}
default: c = esc; break;
}
@@ -3756,8 +3610,10 @@ el_val_t json_get_raw(el_val_t json_str, el_val_t key) {
const char* k = EL_CSTR(key);
const char* p = json_find_key(json, k);
/* Clear fs_read binary-length hint — result is a fresh null-terminated
* string, not the raw file bytes, so Content-Length must use strlen. */
* string, not the raw file bytes, so Content-Length must use strlen.
* (Kept although the pointer pairing now makes this redundant.) */
_tl_fs_read_len = 0;
_tl_fs_read_buf = NULL;
if (!p) return el_wrap_str(el_strdup(""));
const char* end = json_skip_value(p);
size_t n = (size_t)(end - p);
@@ -6228,13 +6084,6 @@ void el_cgi_init(el_val_t name, el_val_t dharma_id, el_val_t principal,
#define ENGRAM_SUPPRESSION_BREAKTHROUGH 5
#define ENGRAM_BREAKTHROUGH_WEIGHT 0.25
#define ENGRAM_INHIBITION_FACTOR 0.1
/* ENGRAM_WM_CAP: hard global ceiling on nodes holding working_memory_weight
* > 0 at any time. Cowan (2001) puts human WM capacity at ~4 chunks; 24 gives
* the daemon generous headroom while preventing the unbounded growth observed
* in production (wm_active 288-778 per heartbeat "working memory" that is
* really the whole recently-touched graph). Ported from release runtime
* v1.0.0-20260501 Pass 5 on 2026-07-15 self-review. */
#define ENGRAM_WM_CAP 24
/* ── Layered consciousness architecture ──────────────────────────────────────
*
@@ -7013,75 +6862,116 @@ static int istr_contains(const char* hay, const char* needle) {
return 0;
}
/* ── Tokenized query matching ───────────────────────────────────────────
* The engram query surface (search / activate / goal-bias) historically
* matched the ENTIRE raw query string as a single case-insensitive
* substring via istr_contains(field, q). That is Ctrl-F, not search:
* a multi-word query like "windows msi signing" only matched a node whose
* text contained that exact contiguous run, so real multi-word queries
* returned zero. istr_contains stays as the per-TOKEN primitive; these
* helpers split the query on whitespace and match ANY token, then rank by
* how many DISTINCT tokens a node covers. Single-token queries are a strict
* special case (score is 0 or 1) so single-word callers never regress. */
#define ENGRAM_MAX_QTOKENS 32
#define ENGRAM_QTOK_LEN 256
/* ---- SPEC-SEARCH-UPGRADE-OURS-2026-07-14: ranked search (BM25 + recency) ----
* Replaces first-N-in-storage-order substring matching (measured 13% hit@5 on
* the 15-query pinned eval; ranked model measured 93% offline). Deterministic,
* local, transparent no model call on the hot path. Multi-word queries score
* per-token (rare+concentrated terms weigh most); ties break newest-first so
* fresh memories stop losing to storage order. The transparent-layer identity
* filter is preserved unchanged: hidden self layers stay invisible here and
* surface only via engram_activate the legitimate path. */
/* Split q on whitespace into up to ENGRAM_MAX_QTOKENS distinct
* (case-insensitive) tokens. Returns the token count. Over-long tokens are
* truncated to ENGRAM_QTOK_LEN-1; over-count tokens are ignored. */
static int engram_tokenize_query(const char* q,
char toks[][ENGRAM_QTOK_LEN], int maxtok) {
#define ENGRAM_BM25_MAX_QTOK 16
#define ENGRAM_BM25_TOKLEN 48
static int engram_tok_next(const char** ps, char* out, int cap) {
const char* s = *ps;
while (*s && !isalnum((unsigned char)*s)) s++;
if (!*s) { *ps = s; return 0; }
int n = 0;
if (!q) return 0;
const char* p = q;
while (*p && n < maxtok) {
while (*p && isspace((unsigned char)*p)) p++;
if (!*p) break;
char buf[ENGRAM_QTOK_LEN];
size_t tl = 0;
while (*p && !isspace((unsigned char)*p)) {
if (tl < sizeof(buf) - 1) buf[tl++] = *p;
p++;
}
buf[tl] = '\0';
if (tl == 0) continue;
int dup = 0;
for (int s = 0; s < n; s++) {
if (strcasecmp(toks[s], buf) == 0) { dup = 1; break; }
}
if (dup) continue;
memcpy(toks[n], buf, tl + 1);
n++;
while (*s && isalnum((unsigned char)*s)) {
if (n < cap - 1) out[n++] = (char)tolower((unsigned char)*s);
s++;
}
return n;
out[n] = 0; *ps = s; return 1;
}
/* Count how many of the ntok distinct query tokens appear (case-insensitive)
* in the node's content, label, or tags. 0 == no match. */
static int engram_node_match_score(const EngramNode* n,
char toks[][ENGRAM_QTOK_LEN], int ntok) {
int score = 0;
for (int t = 0; t < ntok; t++) {
if (istr_contains(n->content, toks[t]) ||
istr_contains(n->label, toks[t]) ||
istr_contains(n->tags, toks[t]))
score++;
static void engram_field_stats(const char* field,
char qtok[][ENGRAM_BM25_TOKLEN], int nq,
int64_t* tf, int64_t* doclen) {
if (!field) return;
char buf[ENGRAM_BM25_TOKLEN];
const char* p = field;
while (engram_tok_next(&p, buf, sizeof buf)) {
(*doclen)++;
for (int t = 0; t < nq; t++)
if (strcmp(buf, qtok[t]) == 0) tf[t]++;
}
return score;
}
/* Rank entry: distinct-token match count (primary, desc) then salience
* (tiebreak, desc). */
typedef struct { int64_t idx; int score; double salience; } EngramRankEntry;
static int engram_rank_cmp(const void* a, const void* b) {
const EngramRankEntry* ea = (const EngramRankEntry*)a;
const EngramRankEntry* eb = (const EngramRankEntry*)b;
if (ea->score != eb->score) return eb->score - ea->score; /* desc */
if (ea->salience < eb->salience) return 1;
if (ea->salience > eb->salience) return -1;
typedef struct { double score; int64_t created; int64_t idx; } EngramHit;
static int engram_hit_cmp(const void* a, const void* b) {
const EngramHit* x = (const EngramHit*)a;
const EngramHit* y = (const EngramHit*)b;
if (x->score != y->score) return (x->score < y->score) ? 1 : -1;
if (x->created != y->created) return (x->created < y->created) ? 1 : -1;
return 0;
}
/* Scores every visible node against the query; writes ranked hits into `out`
* (caller allocates g->node_count entries). Returns min(hits, lim). */
static int64_t engram_search_ranked(EngramStore* g, const char* q, int64_t lim,
EngramHit* out) {
char qtok[ENGRAM_BM25_MAX_QTOK][ENGRAM_BM25_TOKLEN];
int nq = 0;
{
const char* p = q; char buf[ENGRAM_BM25_TOKLEN];
while (nq < ENGRAM_BM25_MAX_QTOK && engram_tok_next(&p, buf, sizeof buf)) {
int dup = 0;
for (int t = 0; t < nq; t++)
if (strcmp(qtok[t], buf) == 0) { dup = 1; break; }
if (!dup) { strcpy(qtok[nq], buf); nq++; }
}
}
if (nq == 0) return 0;
int64_t N = g->node_count;
int64_t* tfm = (int64_t*)calloc((size_t)(N * nq), sizeof(int64_t));
int64_t* dlen = (int64_t*)calloc((size_t)N, sizeof(int64_t));
if (!tfm || !dlen) { free(tfm); free(dlen); return 0; }
int64_t df[ENGRAM_BM25_MAX_QTOK] = {0};
double total_len = 0.0; int64_t live = 0;
for (int64_t i = 0; i < N; i++) {
EngramNode* n = &g->nodes[i];
if (engram_layer_is_transparent(n->layer_id)) continue;
live++;
int64_t* tf = &tfm[i * nq];
engram_field_stats(n->content, qtok, nq, tf, &dlen[i]);
engram_field_stats(n->label, qtok, nq, tf, &dlen[i]);
engram_field_stats(n->tags, qtok, nq, tf, &dlen[i]);
total_len += (double)dlen[i];
for (int t = 0; t < nq; t++) if (tf[t] > 0) df[t]++;
}
double avg = (live > 0) ? total_len / (double)live : 1.0;
if (avg <= 0.0) avg = 1.0;
const double k1 = 1.2, b = 0.75;
int64_t nhits = 0;
for (int64_t i = 0; i < N; i++) {
EngramNode* n = &g->nodes[i];
if (engram_layer_is_transparent(n->layer_id)) continue;
int64_t* tf = &tfm[i * nq];
double s = 0.0;
for (int t = 0; t < nq; t++) {
if (tf[t] == 0) continue;
double idf = log(((double)live - (double)df[t] + 0.5) /
((double)df[t] + 0.5) + 1.0);
double tfd = (double)tf[t];
s += idf * (tfd * (k1 + 1.0)) /
(tfd + k1 * (1.0 - b + b * (double)dlen[i] / avg));
}
if (s > 0.0) {
out[nhits].score = s;
out[nhits].created = n->created_at;
out[nhits].idx = i;
nhits++;
}
}
free(tfm); free(dlen);
qsort(out, (size_t)nhits, sizeof(EngramHit), engram_hit_cmp);
return (nhits < lim) ? nhits : lim;
}
el_val_t engram_search(el_val_t query, el_val_t limit) {
EngramStore* g = engram_get();
const char* q = EL_CSTR(query);
@@ -7089,33 +6979,12 @@ el_val_t engram_search(el_val_t query, el_val_t limit) {
if (lim <= 0) lim = 100;
el_val_t lst = el_list_empty();
if (!q || !*q) return lst;
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
int ntok = engram_tokenize_query(q, toks, ENGRAM_MAX_QTOKENS);
if (ntok == 0) return lst;
EngramRankEntry* hits = malloc((size_t)g->node_count * sizeof(EngramRankEntry));
if (g->node_count == 0) return lst;
EngramHit* hits = (EngramHit*)malloc((size_t)g->node_count * sizeof(EngramHit));
if (!hits) return lst;
int64_t nhits = 0;
for (int64_t i = 0; i < g->node_count; i++) {
EngramNode* n = &g->nodes[i];
/* Filter transparent layers: nodes whose layer is `transparent=1`
* shape output but are invisible to introspection ("what do you
* know about yourself"). They still surface via engram_activate
* + engram_compile_layered_json that's the legitimate path. */
if (engram_layer_is_transparent(n->layer_id)) continue;
int sc = engram_node_match_score(n, toks, ntok);
if (sc > 0) {
hits[nhits].idx = i;
hits[nhits].score = sc;
hits[nhits].salience = n->salience;
nhits++;
}
}
/* Rank by distinct tokens matched (desc) then salience (desc), then cap. */
qsort(hits, (size_t)nhits, sizeof(EngramRankEntry), engram_rank_cmp);
int64_t end = nhits < lim ? nhits : lim;
for (int64_t k = 0; k < end; k++) {
lst = el_list_append(lst, engram_node_to_map(&g->nodes[hits[k].idx]));
}
int64_t k = engram_search_ranked(g, q, lim, hits);
for (int64_t i = 0; i < k; i++)
lst = el_list_append(lst, engram_node_to_map(&g->nodes[hits[i].idx]));
free(hits);
return lst;
}
@@ -7393,14 +7262,10 @@ static double engram_temporal_proximity_bonus(int64_t node_created,
static double engram_goal_bias(const EngramNode* n, const char* query) {
if (!query || !*query) return 1.0;
double bias = 1.0;
/* Direct lexical overlap, graded by token coverage: a node covering all
* query tokens gets the full +0.5; partial coverage gets a proportional
* share. Single-token queries full +0.5 on match, identical to before. */
{
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
int ntok = engram_tokenize_query(query, toks, ENGRAM_MAX_QTOKENS);
int sc = engram_node_match_score(n, toks, ntok);
if (sc > 0 && ntok > 0) bias += 0.5 * ((double)sc / (double)ntok);
/* Direct lexical overlap: node content/label/tags share text with query. */
if (istr_contains(n->content, query) || istr_contains(n->label, query) ||
istr_contains(n->tags, query)) {
bias += 0.5;
}
/* Node-type resonance with query intent. */
int technical_query = istr_contains(query, "code") ||
@@ -7439,48 +7304,6 @@ static double engram_goal_bias(const EngramNode* n, const char* query) {
return bias;
}
/* eg_cmp_double_desc — qsort comparator, descending doubles. */
static int eg_cmp_double_desc(const void* a, const void* b) {
double da = *(const double*)a, db = *(const double*)b;
if (da < db) return 1;
if (da > db) return -1;
return 0;
}
/* eg_enforce_wm_cap_global — clamp the store-wide working-memory population
* to ENGRAM_WM_CAP, keeping the top-K by current weight. Runs at every point
* that materializes WM: post-activation persist and snapshot load/merge.
* (Ported from release runtime v1.0.0-20260501 Pass 5, 2026-07-15.) */
static void eg_enforce_wm_cap_global(EngramStore* g) {
int64_t wm_count = 0;
for (int64_t i = 0; i < g->node_count; i++) {
if (g->nodes[i].working_memory_weight > 0.0) wm_count++;
}
if (wm_count <= ENGRAM_WM_CAP) return;
double* vals = malloc((size_t)wm_count * sizeof(double));
if (!vals) return; /* OOM: over cap this call, no corruption */
int64_t vi = 0;
for (int64_t i = 0; i < g->node_count; i++) {
if (g->nodes[i].working_memory_weight > 0.0)
vals[vi++] = g->nodes[i].working_memory_weight;
}
qsort(vals, (size_t)wm_count, sizeof(double), eg_cmp_double_desc);
double cutoff = vals[ENGRAM_WM_CAP - 1];
free(vals);
int64_t above = 0;
for (int64_t i = 0; i < g->node_count; i++) {
if (g->nodes[i].working_memory_weight > cutoff) above++;
}
int64_t slots_at_cutoff = ENGRAM_WM_CAP - above;
for (int64_t i = 0; i < g->node_count; i++) {
EngramNode* n = &g->nodes[i];
if (n->working_memory_weight <= 0.0) continue;
if (n->working_memory_weight > cutoff) continue;
if (slots_at_cutoff > 0) { slots_at_cutoff--; continue; }
n->working_memory_weight = 0.0; /* evict: over global cap */
}
}
el_val_t engram_activate(el_val_t query, el_val_t depth) {
EngramStore* g = engram_get();
const char* q = EL_CSTR(query);
@@ -7508,21 +7331,14 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
if (!seeds) {
free(best_bg); free(best_hops); free(reached); return out;
}
/* Tokenize once: a node seeds if it matches ANY query token, and its seed
* activation is scaled by token coverage (fraction of distinct query
* tokens it contains) so a node matching all words seeds more strongly
* than one matching a single word. Single-word queries coverage 1.0,
* identical to the prior whole-query behavior. */
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
int ntok = engram_tokenize_query(q, toks, ENGRAM_MAX_QTOKENS);
for (int64_t i = 0; i < g->node_count; i++) {
EngramNode* n = &g->nodes[i];
int sc = engram_node_match_score(n, toks, ntok);
if (sc > 0) {
if (istr_contains(n->content, q) ||
istr_contains(n->label, q) ||
istr_contains(n->tags, q)) {
double tdecay = engram_temporal_decay(n, now_ms);
double dampen = engram_activation_dampen(n);
double cover = ntok > 0 ? (double)sc / (double)ntok : 1.0;
double act = n->salience * tdecay * dampen * cover;
double act = n->salience * tdecay * dampen;
seeds[seed_count].idx = i;
seeds[seed_count].act = act;
seeds[seed_count].created_at = n->created_at;
@@ -7686,12 +7502,6 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
g->nodes[i].working_memory_weight = wm_weights[i];
}
/* Global WM cap: keep only the top ENGRAM_WM_CAP by weight across the
* whole store (see eg_enforce_wm_cap_global). Without this, repeated
* activation calls accumulate hundreds of "promoted" nodes and WM stops
* meaning anything (production heartbeats showed wm_active up to 778). */
eg_enforce_wm_cap_global(g);
/* ── Collect all background-activated nodes for the return value ────
* Callers see both layers. Context compilation uses only promoted nodes
* (working_memory_weight > 0). Sort: promoted first by wm_weight desc,
@@ -8069,9 +7879,6 @@ el_val_t engram_load(el_val_t path) {
}
}
}
/* WM cap discipline applies to every entry point that materializes WM,
* including snapshot restore (see eg_enforce_wm_cap_global). */
eg_enforce_wm_cap_global(g);
free(data);
return 1;
}
@@ -8092,72 +7899,42 @@ el_val_t engram_get_node_json(el_val_t id) {
return el_wrap_str(jb_finish(&b));
}
/* engram_get_node_by_label — find the first node whose label field exactly
* matches the given string. Returns the node as a JSON object string, or "{}"
* if no match is found.
*
* Exact match (strcmp, not substring) because labels like "conv:history"
* must not collide with nodes whose content contains that substring.
*
* Ported from the release runtime 2026-07-16 self-review: chat.el has called
* this since 2026-07-01 but the function only existed in
* releases/v1.0.0-20260501/el_runtime.c the soul daemon (which builds
* against THIS runtime) failed to compile once clang made implicit
* declarations an error. */
/* Look up a node by exact label; returns its JSON or {}. Ported from the
* v1.0.0 release runtime needed by soul.el session continuity
* (conv_history_load / session_summary_write / emit_session_start_event). */
el_val_t engram_get_node_by_label(el_val_t label) {
const char* lbl = EL_CSTR(label);
if (!lbl || !*lbl) return el_wrap_str(el_strdup("{}"));
if (!lbl || !*lbl) return el_wrap_str(el_strdup(""));
EngramStore* g = engram_get();
for (int64_t i = 0; i < g->node_count; i++) {
EngramNode* n = &g->nodes[i];
if (n->label && strcmp(n->label, lbl) == 0) {
JsonBuf b; jb_init(&b);
engram_emit_node_json(&b, n);
return el_wrap_str(jb_finish(&b));
return el_wrap_str(b.buf);
}
}
return el_wrap_str(el_strdup("{}"));
return el_wrap_str(el_strdup(""));
}
el_val_t engram_search_json(el_val_t query, el_val_t limit) {
/* SPEC-SEARCH-UPGRADE 2026-07-14: same ranked BM25+recency core as
* engram_search; transparent-layer identity filter enforced inside it. */
EngramStore* g = engram_get();
const char* q = EL_CSTR(query);
int64_t lim = (int64_t)limit;
if (lim <= 0) lim = 100;
JsonBuf b; jb_init(&b);
jb_putc(&b, '[');
int first = 1;
if (q && *q) {
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
int ntok = engram_tokenize_query(q, toks, ENGRAM_MAX_QTOKENS);
if (ntok > 0) {
EngramRankEntry* hits =
malloc((size_t)g->node_count * sizeof(EngramRankEntry));
if (hits) {
int64_t nhits = 0;
for (int64_t i = 0; i < g->node_count; i++) {
EngramNode* n = &g->nodes[i];
/* Filter transparent layers — same as engram_search. */
if (engram_layer_is_transparent(n->layer_id)) continue;
int sc = engram_node_match_score(n, toks, ntok);
if (sc > 0) {
hits[nhits].idx = i;
hits[nhits].score = sc;
hits[nhits].salience = n->salience;
nhits++;
}
}
/* Rank by distinct tokens matched (desc) then salience (desc). */
qsort(hits, (size_t)nhits, sizeof(EngramRankEntry),
engram_rank_cmp);
int64_t end = nhits < lim ? nhits : lim;
for (int64_t k = 0; k < end; k++) {
if (!first) jb_putc(&b, ',');
engram_emit_node_json(&b, &g->nodes[hits[k].idx]);
first = 0;
}
free(hits);
if (q && *q && g->node_count > 0) {
EngramHit* hits = (EngramHit*)malloc((size_t)g->node_count * sizeof(EngramHit));
if (hits) {
int64_t k = engram_search_ranked(g, q, lim, hits);
for (int64_t i = 0; i < k; i++) {
if (i) jb_putc(&b, ',');
engram_emit_node_json(&b, &g->nodes[hits[i].idx]);
}
free(hits);
}
}
jb_putc(&b, ']');
@@ -8676,9 +8453,6 @@ el_val_t engram_load_merge(el_val_t path) {
}
}
/* Merged nodes can carry snapshot WM weights too — hold the cap here as
* well (see eg_enforce_wm_cap_global). */
eg_enforce_wm_cap_global(g);
free(data);
return (el_val_t)added_nodes;
}
-1
View File
@@ -632,7 +632,6 @@ el_val_t engram_load(el_val_t path);
* can pass results straight through without round-tripping ElList/ElMap
* through json_stringify. */
el_val_t engram_get_node_json(el_val_t id);
el_val_t engram_get_node_by_label(el_val_t label);
el_val_t engram_search_json(el_val_t query, el_val_t limit);
el_val_t engram_scan_nodes_json(el_val_t limit, el_val_t offset);
el_val_t engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, el_val_t offset);
+1 -25
View File
@@ -23,29 +23,10 @@ fn tok_at(tokens: [Any], pos: Int) -> Map<String, Any> {
}
fn tok_kind(tokens: [Any], pos: Int) -> String {
// Out-of-range reads must report the Eof sentinel so every `== "Eof"`
// termination guard in the parser fires. Without this, reading past the
// single trailing Eof token returns runtime null (el_list_get OOB -> 0),
// which matches no delimiter, letting inner parse loops append AST nodes
// forever on malformed input -> unbounded allocation -> OOM.
let n: Int = native_list_len(tokens) / 2
if pos < 0 {
return "Eof"
}
if pos >= n {
return "Eof"
}
native_list_get(tokens, pos * 2)
}
fn tok_value(tokens: [Any], pos: Int) -> String {
let n: Int = native_list_len(tokens) / 2
if pos < 0 {
return ""
}
if pos >= n {
return ""
}
native_list_get(tokens, pos * 2 + 1)
}
@@ -54,12 +35,7 @@ fn expect(tokens: [Any], pos: Int, kind: String) -> Int {
if k == kind {
return pos + 1
}
// On mismatch, error recovery is best-effort. But never step PAST the Eof
// sentinel: once at Eof a mismatch means the input ended early, and
// advancing would run the cursor off the token list.
if k == "Eof" {
return pos
}
// On mismatch just advance; error recovery is best-effort
pos + 1
}
File diff suppressed because it is too large Load Diff
@@ -117,15 +117,6 @@ el_val_t el_min(el_val_t a, el_val_t b);
void el_retain(el_val_t v);
void el_release(el_val_t v);
/* ── Arena scoping ────────────────────────────────────────────────────────────
* el_arena_push() activates the string arena (if not already active) and
* returns a mark; el_arena_pop(mark) frees all strings allocated since that
* mark. Used by codegen for per-function/statement scoping and by long-running
* EL loops (e.g. the soul daemon's awareness tick) to reclaim per-iteration
* allocations. */
el_val_t el_arena_push(void);
el_val_t el_arena_pop(el_val_t mark);
/* ── List ────────────────────────────────────────────────────────────────── */
el_val_t el_list_new(el_val_t count, ...);
@@ -151,7 +142,6 @@ el_val_t http_get_with_headers(el_val_t url, el_val_t headers_map);
el_val_t http_post_with_headers(el_val_t url, el_val_t body, el_val_t headers_map);
el_val_t http_post_form_auth(el_val_t url, el_val_t form_body, el_val_t auth_header);
el_val_t http_delete(el_val_t url);
el_val_t http_delete_json(el_val_t url, el_val_t json_body);
void http_serve(el_val_t port, el_val_t handler);
void http_set_handler(el_val_t name);
@@ -177,11 +167,6 @@ void http_set_handler(el_val_t name);
void http_serve_v2(el_val_t port, el_val_t handler);
void http_set_handler_v2(el_val_t name);
/* Non-blocking variant of http_serve: runs the accept loop in a background
* pthread and returns immediately so the caller can continue (used by the
* soul daemon to run awareness_run() after starting its HTTP API). */
void http_serve_async(el_val_t port, el_val_t handler);
/* Build an HTTP response envelope. `headers_json` should be a JSON object
* literal like `{"WWW-Authenticate":"Basic"}` (or "" / "{}" for none). The
* returned string carries the discriminator `{"el_http_response":1,...}`
@@ -591,7 +576,6 @@ el_val_t engram_list_layers(void);
el_val_t engram_get_node(el_val_t id);
void engram_strengthen(el_val_t node_id);
void engram_forget(el_val_t node_id);
el_val_t engram_prune_telemetry(el_val_t older_than_ms);
el_val_t engram_node_count(void);
el_val_t engram_search(el_val_t query, el_val_t limit);
el_val_t engram_scan_nodes(el_val_t limit, el_val_t offset);
@@ -610,32 +594,12 @@ el_val_t engram_load(el_val_t path);
* can pass results straight through without round-tripping ElList/ElMap
* through json_stringify. */
el_val_t engram_get_node_json(el_val_t id);
el_val_t engram_get_node_by_label(el_val_t label);
el_val_t engram_search_json(el_val_t query, el_val_t limit);
el_val_t engram_scan_nodes_json(el_val_t limit, el_val_t offset);
el_val_t engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, el_val_t offset);
el_val_t engram_neighbors_json(el_val_t node_id, el_val_t max_depth, el_val_t direction);
el_val_t engram_activate_json(el_val_t query, el_val_t depth);
el_val_t engram_stats_json(void);
el_val_t engram_act_stats_json(void);
el_val_t engram_text_health_json(void);
el_val_t engram_cosine_sim(el_val_t id_a, el_val_t id_b);
/* Destructively pop up to `max` newly-formed Hebbian associations as a JSON
* array of {from_id,to_id,weight,hebb}. The learning process (soul daemon) is
* not the process that owns persistence (engram HTTP server); this is how a
* self-formed association crosses that boundary. (2026-08-07 self-review.) */
el_val_t engram_hebb_drain_json(el_val_t max);
/* Document frequency of a term across node labels — term-specificity signal
* for curiosity seed selection. (2026-08-03 self-review.) */
el_val_t engram_label_df(el_val_t term);
/* Best curiosity seed from one node: argmax over idf·position·casing across
* the candidate tokens of its label, falling back to its content when the
* label is a sentinel. Excludes pipe-delimited tabu terms during selection
* and gates candidates to the df band [min_df, max_df]. Returns "" when
* nothing qualifies. (2026-08-13 self-review.) */
el_val_t engram_salient_term(el_val_t node_id, el_val_t max_df,
el_val_t min_df, el_val_t tabu);
el_val_t engram_embed_backfill(el_val_t count);
el_val_t engram_list_layers_json(void);
/* Working memory introspection — count, mean weight, and top-N snapshot.
* Ported from el-compiler/runtime on 2026-06-30 self-review. */