Compare commits

..

26 Commits

Author SHA1 Message Date
will.anderson a118d19393 Merge pull request 'promote dev -> stage: el cluster (#66 engram + #79 truncation fix + release-runtime Windows port)' (#81) from dev into stage
El SDK CI - stage / build-and-test (push) Successful in 7m58s
El SDK Release / build-and-release (pull_request) Successful in 4m16s
2026-07-22 21:20:17 +00:00
will.anderson c6aa1e5c53 Merge pull request 'Land el cluster: #66 engram search + natives, #79 truncation fix, + release-runtime Windows port (reconciled)' (#80) from reconcile/el-cluster-windows-runtime into dev
El SDK CI - dev / build-and-test (push) Successful in 8m3s
El SDK CI - stage / build-and-test (pull_request) Successful in 4m25s
Land el cluster (#66 + #79 + release-runtime Windows-port reconciliation) into dev
2026-07-22 21:06:36 +00:00
will.anderson ff577391f2 reconcile(release-runtime): Windows-port + complete v1.0.0 release runtime so the desktop soul cross-compiles
El SDK CI - dev / build-and-test (pull_request) Successful in 7m7s
The desktop soul (neuron/dist) compiles against the v1.0.0-20260501 release
runtime. After #66 landed the engram natives (tokenized/ranked search,
engram_prune_telemetry) and #79 the durable truncation fix into this runtime,
two gaps remained before it could cross-compile the Windows brain:

1. Windows OS boundary: the release runtime had no Win32 path. Ported the same
   _WIN32-guarded shim the mainline runtime carries (#69): #ifdef _WIN32 ->
   el_platform_win.h (winsock/dlsym/popen + WSAStartup ctor), SOCKET fd guards
   and el_closesocket() at every socket site, CreateProcessA for exec_bg, the
   tm_zone/mingw guard, an el_setsockopt optval wrapper (GCC14), and curl-less
   libcurl stubs. Every change is _WIN32/HAVE_CURL-gated — the POSIX build is
   byte-identical (gcc -fsyntax-only clean; native behaviour unchanged).

2. Header exports: the release el_runtime.h omitted symbols the soul dist calls
   that are defined in this runtime's .c — the http_handler_fn/http_handler4_fn
   typedefs and el_arena_push/pop, engram_prune_telemetry, engram_get_node_by_label.
   Declaration-only, POSIX-neutral; fixes implicit-declaration/unknown-type
   errors under the C11 mingw build.

Result: x86_64-w64-mingw32-gcc compiles el_runtime.c + all 48 soul modules
clean; POSIX gcc -fsyntax-only clean. This is the Windows-port PR the runtime
needed on main (the release-runtime counterpart to #69), landed via stage.
2026-07-22 15:56:08 -05:00
will.anderson ee0d5f9b97 Merge #79: durable HTTP response-truncation fix, both runtimes (via stage) 2026-07-22 15:46:02 -05:00
will.anderson 391bd818ea Merge #66: tokenized+ranked engram lexical search + engram natives (via stage) 2026-07-22 15:45:54 -05:00
will.anderson 43636aed99 runtime: pair fs_read length hint with its buffer in BOTH runtimes — kill response truncation for good
El SDK Release / build-and-release (pull_request) Failing after 7s
The binary-safe fs_read length (_tl_fs_read_len) was consumed by the HTTP
response path for ANY body, even when a handler wrapped a smaller file into a
larger reply. Content-Length then lied AND the send stopped short: the
safety-contact (988) routes returned 178 of 208/218 bytes, cut mid-'set_at' —
unparseable JSON. The desktop app read that as failure. On Windows the shipped
brain is an OLD build without even the per-handler workaround, so EVERY reply
truncated: the app can't read confirmations and refuses the new user.

Durable fix: pair the length hint with the exact buffer pointer it describes
(_tl_fs_read_buf). Apply the raw byte count ONLY when the response IS that
buffer (binary file serving stays correct); every wrapped/enveloped/derived
body is measured with strlen. Reset both at request start and in fs_read /
json_get_raw. This also closes the stale-hint heap over-read (a length larger
than a later body would read past it out the socket) that a plain max() leaves
open — so this class of bug dies on every platform, not just where a handler
happened to be patched.

Applied identically to the mainline runtime (lang/el-compiler/runtime) AND the
frozen release runtime (lang/releases/v1.0.0-20260501) the desktop souls
compile against — the release copy still carried the raw leak, which is why the
Windows brain kept truncating. Same proven approach as PR #78 (Tim Lingo),
extended to cover the release runtime and rebased onto current main.

Both runtimes: gcc -fsyntax-only clean.
2026-07-22 15:04:25 -05:00
will.anderson 2baa0b9a41 Merge pull request 'release: promote stage -> main (ci publish hardening for sdk-release)' (#77) from stage into main
El SDK Release / build-and-release (push) Successful in 7m55s
2026-07-15 21:21:28 +00:00
will.anderson 6a8b2461cd Merge pull request 'release: promote dev -> stage (ci publish hardening for stage/main)' (#76) from dev into stage
El SDK CI - stage / build-and-test (push) Successful in 8m19s
El SDK Release / build-and-release (pull_request) Failing after 13m1s
2026-07-15 21:16:11 +00:00
will.anderson bcb356fe69 Merge pull request 'ci(stage,main): decouple ci-base rebuild, make SDK publish fail loudly' (#75) from hotfix/ci-stage-main-publish-hardening into dev
El SDK CI - stage / build-and-test (pull_request) Successful in 4m27s
El SDK CI - dev / build-and-test (push) Failing after 14m3s
2026-07-15 21:15:27 +00:00
will.anderson dd7827059a ci(stage,main): decouple ci-base rebuild, make SDK publish fail loudly
El SDK CI - dev / build-and-test (pull_request) Failing after 14m30s
Mirror the PR #72 fix (applied to ci-dev.yaml) onto ci-stage.yaml and
sdk-release.yaml. The stage and prod release jobs reported FAILURE even
when the el-runtime-c/-h publish SUCCEEDED, because the ancillary ci-base
Docker rebuild (a CI-cache optimization on the fragile host-mode GCE
runner) reddened the whole job.

- Rebuild ci-base step: continue-on-error: true — never blocks/reddens
  the job; the SDK publish is the deliverable.
- Publish step: set -euo pipefail + empty-key guard + active-account echo
  so a real publish failure still fails loud and is diagnosable.
2026-07-15 16:14:50 -05:00
will.anderson 208e36c899 Merge pull request 'release: promote stage -> main (tokenized search, get_node_by_label, epm fix, win portability)' (#74) from stage into main
El SDK Release / build-and-release (push) Successful in 8m23s
2026-07-15 18:24:39 +00:00
will.anderson b97ce74d1f Merge pull request 'release: promote dev -> stage (tokenized search, get_node_by_label, epm fix)' (#73) from dev into stage
El SDK CI - stage / build-and-test (push) Failing after 8m45s
El SDK Release / build-and-release (pull_request) Successful in 4m1s
2026-07-15 17:20:11 +00:00
will.anderson 155a449c4e Merge pull request 'ci(dev): make SDK publish fail loudly, decouple ci-base rebuild' (#72) from hotfix/ci-dev-publish-hardening into dev
El SDK CI - dev / build-and-test (push) Successful in 8m56s
El SDK CI - stage / build-and-test (pull_request) Successful in 4m10s
2026-07-15 16:34:14 +00:00
will.anderson 4696fd6833 ci(dev): make SDK publish fail loudly, decouple ci-base rebuild
El SDK CI - dev / build-and-test (pull_request) Successful in 8m51s
The dev push build went green-then-red while nothing published: the
Publish step had no set -e, so an auth/upload failure exited 0 (silent
no-publish), while the ci-base rebuild (set -euo pipefail + Docker on the
host-mode runner) hard-failed the job. Add set -euo pipefail + an empty-key
guard + active-account echo to the Publish step so failures surface with a
retrievable log, and mark the ci-base cache rebuild continue-on-error so
the fragile Docker step can never block the actual SDK artifact publish.
2026-07-15 11:33:37 -05:00
will.anderson 581a351fb1 Merge pull request 'integrate: stack PRs #65–#69 (elc OOM guard, tokenized+semantic engram search, get_node_by_label, win portability) for green CI' (#71) from hotfix/stage-elc-engram-integration into dev
El SDK CI - dev / build-and-test (push) Failing after 14m31s
2026-07-15 15:49:57 +00:00
will.anderson 8ce8656de2 epm: declare cross-module callees as extern fn so strict compilers accept generated C
El SDK CI - dev / build-and-test (pull_request) Successful in 7m33s
epm's sibling modules (registry/install/update) call functions defined in other
modules and in the El runtime (config, read_installed, registry_find,
manifest_deps, manifest_name, registry_latest_version, registry_token,
install_vessel, installed_version) without importing them, so elc emits no C
prototype for those calls. gcc<=13 treated the resulting implicit declarations
as warnings; gcc>=14 and clang reject them as hard errors, which is why the
"Build epm" CI step fails and blocks the whole dev/stage pipeline.

Add `extern fn` forward declarations -- El's own separate-compilation mechanism
-- for each cross-module callee at the top of registry/install/update. This
gives elc the correct C prototype in every generated translation unit, so the
calls compile cleanly and still resolve at link time. Simply suppressing
-Wimplicit-function-declaration would be unsafe: an implicit int return
truncates the 64-bit pointer returns of config/registry_find into a latent
crash, so declaring the true signatures is the correct fix. Localized to epm;
touches neither elc nor the runtime.
2026-07-15 10:14:43 -05:00
will.anderson 1e49560f1f Merge remote-tracking branch 'origin/feat/engram-semantic-search' into hotfix/stage-elc-engram-integration
El SDK CI - dev / build-and-test (pull_request) Failing after 14m39s
# Conflicts:
#	lang/el-compiler/runtime/el_runtime.c
2026-07-15 09:33:05 -05:00
will.anderson e8f0b5a9de Merge remote-tracking branch 'origin/fix/engram-lexical-tokenized-search' into hotfix/stage-elc-engram-integration 2026-07-15 09:28:44 -05:00
will.anderson 40287c4cfc Merge remote-tracking branch 'origin/hotfix/win-runtime-portability' into hotfix/stage-elc-engram-integration 2026-07-15 09:28:44 -05:00
will.anderson 0481bea44d Merge remote-tracking branch 'origin/hotfix/runtime-engram-get-node-by-label' into hotfix/stage-elc-engram-integration 2026-07-15 09:28:44 -05:00
will.anderson 9d565ca080 Merge remote-tracking branch 'origin/hotfix/elc-fixes' into hotfix/stage-elc-engram-integration 2026-07-15 09:28:44 -05:00
will.anderson 4773dd0aa2 runtime: make Windows soul reproducible from a clean el checkout
El SDK Release / build-and-release (pull_request) Failing after 16s
Two el_runtime portability defects only ever lived in staged local copies
used to hand-build neuron-ui PR #136's curl-enabled Windows neuron.exe.
gcc 15 promotes both to hard errors, so a clean el checkout cannot rebuild
that soul. Upstream the minimal fixes so the build is reproducible:

- http_serve_async: cast setsockopt optval to (const char*). Win32/mingw
  setsockopt wants const char*, not int*; the cast is a no-op on POSIX and
  matches the four already-cast sites elsewhere in this file.
- engram_save persist path: map fsync -> _commit in the _WIN32-only
  el_platform_win.h (io.h already included). Windows has no fsync(); the
  POSIX path is untouched.
2026-07-15 04:24:08 -05:00
will.anderson 6b9d9e6c4a Add engram_get_node_by_label runtime native to unblock soul link
El SDK Release / build-and-release (pull_request) Failing after 22s
chat.el calls the runtime native engram_get_node_by_label to fetch
well-known nodes (conv:history, session:summary) by stable label rather
than by ID — immune to vector-index drift across restarts. The current
runtime never defined it, so the regenerated dist/soul.c fails to link.

Backport the function verbatim (idiom-adapted to jb_finish) from release
runtime v1.0.0-20260501 and register it as an EL builtin exactly like its
siblings: runtime definition + prototype, __-prefixed seed wrapper +
prototype, and codegen arity entry. No search-site code is touched.
2026-07-15 04:07:33 -05:00
will.anderson b4967af13e feat(engram): semantic search layer via nomic-embed-text (cosine ∪ lexical)
Lexical istr_contains alone can't surface a node whose words don't appear
in the query. This adds an optional dense-vector layer: node content and the
query are embedded through Ollama (nomic-embed-text), and nodes are ranked by
cosine similarity unioned with lexical hits, so a paraphrase query reaches the
right node.

Wired into all three query entry points in el_runtime.c:
  - engram_search_json (HTTP /api/search): collect lexical ∪ semantic
    candidates, score (lexical base 1.0 + cosine; pure-semantic = cosine),
    rank, emit top-N. Stable sort preserves old order when semantic is off.
  - engram_search (internal el_val twin): lexical ∪ semantic union.
  - engram_activate seed loop (HTTP /api/activate): a node seeds if it
    lexically matches OR clears the cosine threshold; pure-semantic seeds
    enter scaled by cosine so paraphrase spreads without overpowering.

Degradable by design: the whole layer is gated on HAVE_CURL plus a one-shot
runtime probe. If curl is compiled out, Ollama is unreachable, or
ENGRAM_SEMANTIC=0, every entry point yields zero semantic signal and callers
fall back byte-for-byte to the pre-existing lexical search.

Node embeddings are cached in process memory keyed by node id with an FNV-1a
content hash for invalidation; the query is embedded once per call — so the
graph is not re-embedded on every query. nomic task prefixes
(search_query:/search_document:) are applied for retrieval separation.

Build steps gain -DHAVE_CURL so the engram artifact compiles the layer in
(-lcurl was already linked). Env: ENGRAM_SEMANTIC, ENGRAM_EMBED_URL,
ENGRAM_EMBED_MODEL, ENGRAM_SEMANTIC_MIN (cosine threshold, default 0.6).
2026-07-14 18:48:16 -05:00
will.anderson 2b2a1246e7 Merge pull request 'runtime: fix the memory-leak + write-corruption pair in el_runtime.c' (#64) from hotfix/el-runtime-leak-and-persist into main
El SDK Release / build-and-release (push) Failing after 10m52s
2026-07-13 21:23:31 +00:00
will.anderson 5c41c66a0f Merge pull request 'fix(windows): guard el_mem_check with _WIN32 — rusage is POSIX-only' (#60) from fix/windows-rusage-guard into stage
El SDK CI - stage / build-and-test (push) Failing after 13m21s
fix(windows): guard el_mem_check with _WIN32 — rusage is POSIX-only
2026-06-25 16:48:13 +00:00
48 changed files with 880 additions and 822776 deletions
+15
View File
@@ -214,9 +214,18 @@ jobs:
env:
GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
run: |
# Fail loudly: previously this step had no `set -e`, so an auth or
# upload failure was swallowed (step exited 0 on the trailing echo)
# and the SDK silently never published. Surface failures now.
set -euo pipefail
if [ -z "${GCP_SA_KEY:-}" ]; then
echo "FATAL: GCP_SA_KEY secret is empty — cannot authenticate to publish" >&2
exit 1
fi
echo "${GCP_SA_KEY}" > /tmp/gcp-key.json
gcloud auth activate-service-account --key-file=/tmp/gcp-key.json
gcloud config set project neuron-785695
echo "Publishing as active account: $(gcloud config get-value account 2>/dev/null)"
VERSION="${GITHUB_SHA:0:8}"
@@ -268,6 +277,12 @@ jobs:
# Patches ci-base:dev in-place: pulls the existing image (which has all
# system deps — Node, Go, gcloud, Docker CLI, etc.) and overlays the freshly
# built El SDK on top. Keeps the full ci-base rebuild fast and incremental.
#
# continue-on-error: this is a CI-cache optimization, NOT the release
# artifact. It runs Docker (pull/build/push ~600MB) on the host-mode GCE
# runner where DinD/Docker availability is fragile. A failure here must
# never block or redden the job — the SDK publish above is the deliverable.
continue-on-error: true
if: github.event_name == 'push'
env:
GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
+15
View File
@@ -212,12 +212,21 @@ jobs:
env:
GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
run: |
# Fail loudly: previously this step had no `set -e`, so an auth or
# upload failure was swallowed (step exited 0 on the trailing echo)
# and the SDK silently never published. Surface failures now.
set -euo pipefail
if [ -z "${GCP_SA_KEY:-}" ]; then
echo "FATAL: GCP_SA_KEY secret is empty — cannot authenticate to publish" >&2
exit 1
fi
echo "${GCP_SA_KEY}" > /tmp/gcp-key.json
apt-get install -y -qq apt-transport-https ca-certificates curl
echo "deb [trusted=yes] https://packages.cloud.google.com/apt cloud-sdk main" > /etc/apt/sources.list.d/google-cloud-sdk.list
apt-get update -qq && apt-get install -y google-cloud-cli
gcloud auth activate-service-account --key-file=/tmp/gcp-key.json
gcloud config set project neuron-785695
echo "Publishing as active account: $(gcloud config get-value account 2>/dev/null)"
VERSION="${GITHUB_SHA:0:8}"
@@ -253,6 +262,12 @@ jobs:
# Patches ci-base:stage in-place: pulls the existing image (which has all
# system deps — Node, Go, gcloud, Docker CLI, etc.) and overlays the freshly
# built El SDK on top. Keeps the full ci-base rebuild fast and incremental.
#
# continue-on-error: this is a CI-cache optimization, NOT the release
# artifact. It runs Docker (pull/build/push ~600MB) on the host-mode GCE
# runner where DinD/Docker availability is fragile. A failure here must
# never block or redden the job — the SDK publish above is the deliverable.
continue-on-error: true
if: github.event_name == 'push'
env:
GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
+15
View File
@@ -288,12 +288,21 @@ jobs:
env:
GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
run: |
# Fail loudly: previously this step had no `set -e`, so an auth or
# upload failure was swallowed (step exited 0 on the trailing echo)
# and the SDK silently never published. Surface failures now.
set -euo pipefail
if [ -z "${GCP_SA_KEY:-}" ]; then
echo "FATAL: GCP_SA_KEY secret is empty — cannot authenticate to publish" >&2
exit 1
fi
echo "${GCP_SA_KEY}" > /tmp/gcp-key.json
apt-get install -y -qq apt-transport-https ca-certificates curl
echo "deb [trusted=yes] https://packages.cloud.google.com/apt cloud-sdk main" > /etc/apt/sources.list.d/google-cloud-sdk.list
apt-get update -qq && apt-get install -y google-cloud-cli
gcloud auth activate-service-account --key-file=/tmp/gcp-key.json
gcloud config set project neuron-785695
echo "Publishing as active account: $(gcloud config get-value account 2>/dev/null)"
VERSION="${GITHUB_SHA:0:8}"
@@ -345,6 +354,12 @@ jobs:
# Patches ci-base:latest in-place: pulls the existing image (which has all
# system deps — Node, Go, gcloud, Docker CLI, etc.) and overlays the freshly
# built El SDK on top. Keeps the full ci-base rebuild fast and incremental.
#
# continue-on-error: this is a CI-cache optimization, NOT the release
# artifact. It runs Docker (pull/build/push ~600MB) on the host-mode GCE
# runner where DinD/Docker availability is fragile. A failure here must
# never block or redden the job — the SDK publish above is the deliverable.
continue-on-error: true
if: github.event_name == 'push'
env:
GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
-65
View File
@@ -1,65 +0,0 @@
# ELP language consolidation — full-lexicon backfill (stage)
Branch: `stage-elp-lang-consolidation` (stage-bound; NOT the live soul :8742).
Consolidates scattered Python language-realizer work (`~/Desktop/lang-realizers`,
`~/Desktop/lang-poetry-experiment`, `~/semitic_engine`) into the ELP `.el`
structure, generating **full lexicons** (complete UniMorph + kaikki.org
Wiktionary — real gender, real inflections) instead of the demo/curated subsets
the prototypes shipped.
## ELP before this branch
- 18 classical/ancient languages fully done (vocab + morphology + tests):
akk ang cop egy enm fro gez goh got grc non peo pi sa sga sux txb uga.
- 11 modern/classical languages had `morphology-<code>.el` in the build manifest
but **no vocabulary and no lang_profile**: es fr de ja ar he hi ru fi sw la.
- The ES port (`stage-elp-es-port`) had a *demo-scale* vocabulary-es.el (~350
entries, s-expr form).
## Landed on this branch (full-lexicon seed-fn format, matching the 18 ancients)
Vocabulary schema per row: `[lemma, pos, form0, form1, form2, en_gloss, hint]`.
Files are ELP runtime **seed data** (loaded via the Engram at runtime), so — like
all 18 classical `vocabulary-*.el` — they are intentionally NOT in the build
manifest. Syntax validated: the chunked `fn vocab_<code>_seed_pN` format
compiles cleanly to C via `elc` (correct UTF-8).
| code | in-ELP-morph? | vocab entries | verbs | nouns | adjs | profile |
|------|---------------|--------------:|------:|------:|-----:|---------|
| es | yes | 72,032 | 6,695 | 48,353 | 16,984 | yes |
| fr | yes | 130,517 | 7,534 | 77,344 | 45,639 | yes |
| de | yes | 144,692 | 6,661 | 133,162 | 4,869 | yes |
| la | yes | 22,590 | 82 | 13,436 | 9,072 | yes |
| it | no (bonus) | 193,675 | 10,008 | 109,459 | 74,208 | yes |
| pt | no (bonus) | 115,772 | 4,001 | 72,073 | 39,698 | yes |
| ro | no (bonus) | 86,504 | 1,216 | 65,915 | 19,373 | yes |
| ca | no (bonus) | 47,112 | 1,547 | 28,830 | 16,735 | yes |
|**total**| |**812,894** | | | | |
Generators (reproducible): `elp/tests/lang-gen/gen_elp_seed_full.py` (Romance),
`gen_elp_seed_de_la.py` (German declension + Latin case-paradigm mapping). They
read the pre-built morph caches in `~/Desktop/lang-realizers/data/` (UniMorph +
kaikki), which are too large to commit.
## Remaining (honest)
Of the 11 ELP backfill targets, 4 are done (es fr de la). The other 7 have **no
full-lexicon engine** yet — cannot be generated honestly without engine work:
- **ru**: only a 110-entry curated Slavic subset exists; full `rus.unimorph`
present but no `morphology_ru_full` productive loader. Needs a full Russian
morphology module (like the Romance ones) before vocab generation.
- **ja / ko / zh**: validated demo engines (~66-104 hardcoded words) in
`lang-poetry-experiment`, Python only. Agglutinative (ja/ko) + isolating (zh)
need `.el` engine ports + full-lexicon wiring (ja: jpn_unimorph; zh: CC-CEDICT).
- **ar / he (Semitic)**: template engines (16 AR / 8 HE patterns, ~6 roots) in
`~/semitic_engine`, Python only. Root-and-pattern; full UniMorph ara/heb
present but used only for validation. Needs productive root lexicon + `.el` port.
- **hi (Hindi), fi (Finnish), sw (Swahili)**: `morphology-<code>.el` exists in
ELP but there is NO scattered prototype and NO downloaded data for these —
full-lexicon collection (UniMorph/kaikki) + generator still to do.
De/nl/sv Germanic and it/ro/ca/pt Romance verb coverage note: German verbs here
are the ~6.6k caches carry; the it/ro/ca/pt bonus languages have full vocab but
**no `morphology-<code>.el` in ELP yet** (Python realizer exists; `.el` port is
the remaining engine work).
Construction coverage (separate from lexicon): French realizer was ~55%,
Semitic ~3% in the prototypes — full construction coverage remains its own task.
-72
View File
@@ -1,72 +0,0 @@
;;; lang_profile_ca.el — Catalan language profile for ELP.
;;; Mirrors lang_profile_it / _es / _pt; keys the realizer's construction switches.
;;; Catalan is the CLOSEST Romance sibling to the shared engine (~85% conceptual
;;; reuse). The deltas: PRONOMS FEBLES with four position allomorphs, l'-elision,
;;; del/al/pel contractions, the periphrastic preterite (vaig+INF), and NO
;;; essere/avere split (perfect aux is always HAVER; ser/estar is only the copula).
(lang_profile_ca
(language "Catalan")
(iso639 "ca")
(family "Romance")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop yes) ; null subjects default; overt pronoun = emphatic
(obligatory-subject no)
(grammatical-gender yes) ; m/f; full NP agreement (art + adj + participle)
(do-support no)
(subject-aux-inversion no) ; yes/no Q = declarative order + '?'; no inversion
(article-selection "el/la/l'/els/les ; un/una/uns/unes") ; l'-ELISION:
; el/la -> l' before vowel or (silent) h, glued to
; the next word (l'home, l'illa); de -> d' before vowel
(article-drives-contraction yes) ; article choice feeds prep+article contraction
(adjective-position "postnominal-default + small prenominal class") ; bo/bon,
; mal, gran, nou, vell, primer, molt... prenominal
(question-punct plain) ; ? and ! only (no inverted ¿ ¡)
;; ── MANDATORY prep+article contractions ────────────────────────────────
(contractions ((de el del) (de els dels)
(a el al) (a els als)
(per el pel) (per els pels)))
(contraction-mandatory yes) ; *de el -> del obligatory
(contraction-blocked-before-elision yes) ; de l'home / a l'home (NO *del home)
;; ── clitic system: PRONOMS FEBLES (the headline delta) ──────────────────
(clitics yes)
(clitic-allomorphy four-position) ; per pronoun, form varies by position+onset:
; reinforced (em, et, el) proclitic before a consonant
; elided (m', t', l', n') proclitic before a vowel/h
; full (-me, -lo, -li) enclitic after a consonant/-r
; reduced ('m, 't, 'l, 'ns) enclitic after a vowel
(clitic-placement ((finite proclitic) ; el veig, no m'ho dóna
(imperative-affirmative enclitic) ; dóna'm, digues-me
(imperative-negative present-subjunctive) ; no parlis (delta)
(infinitive enclitic) ; ajudar-me, veure'l
(gerund enclitic))) ; fent-ho
(clitic-combination ((me el "me'l") (te el "te'l") (se el "se'l")
(me la "me la") (me en "me'n")
(li el "l'hi") (li en "n'hi"))) ; dative+accusative clusters
(clitic-particles (hi en ho)) ; locative hi, partitive/genitive en, neuter ho
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "person+number (6-way)")
(tenses (present imperfet preterit-simple perifrastic-preterit futur
condicional subjuntiu-present subjuntiu-imperfet imperatiu))
(periphrastic-preterite "vaig/vas/va/vam/vau/van + INFINITIVE") ; << hallmark CA
; (vaig cantar = 'I sang'); coexists w/ synthetic pret.
(compound-past "pretèrit perfet = haver(present) + participle")
(perfect-aux "HAVER only") ; << NO essere/avere split (simpler than IT)
(participle-agreement ((haver preceding-acc-clitic))) ; les he vistes; else invariable
(progressive-aux "estar + gerundi")
(copula "ser / estar") ; ser: identity/essential/origin; estar:
; location + transient state (estic cansat, és a casa)
(passive-aux "ser (+ per-agent)")
(future inflectional) ; cantaré, serà
(comparative "més/menys ADJ que")
;; ── SACRED safety bar (shared with es/pt/it/en) ────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
(negation "no (preverbal) + optional 'pas' + concord") ; no...res/
; ningú/mai/cap/gens/enlloc
(negative-concord yes) ; preverbal negative subject (ningú) keeps 'no'
(neg-reinforcer pas)) ; optional (no ho faré pas)
-41
View File
@@ -1,41 +0,0 @@
;;; lang_profile_de.el — German language profile for ELP.
;;; Mirrors lang_profile_en / lang_profile_es. Keys the realizer's construction
;;; switches. German is the largest Germanic delta from the EN engine: V2 word
;;; order, four morphological cases, and separable-prefix verbs.
(lang_profile_de
(language "German")
(iso639 "de")
(family "Germanic")
(neighbor-base "en") ; realized by extending the English (Germanic) engine
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop no) ; obligatory subject in finite clauses
(obligatory-subject yes)
(grammatical-gender (m f n)) ; three genders; drives article + adj declension
(case-system (nom acc dat gen)) ; four cases on articles/adjs/nouns
(word-order V2) ; finite verb 2nd in main clause
(subordinate-order verb-final) ; "..., dass er den Hund SIEHT."
(separable-verbs yes) ; aufstehen -> "steht ... auf"; ppart "aufgestanden"
(do-support no) ; German negates/questions the finite verb directly
(subject-verb-inversion yes) ; yes/no Q fronts finite verb; wh-Q fills Vorfeld
(article-selection "der/die/das + ein/kein") ; declined by case x gender x number
(adjective-position prenominal)
(adjective-declension (strong weak mixed)) ; chosen by the determiner type
(noun-capitalization yes)
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "person-and-number") ; full present/past paradigm
(auxiliary-order (modal tense-aux perfect passive main))
(perfect-aux (haben sein)) ; sein for intransitive motion/change verbs
(passive-aux "werden")
(future "werden + infinitive")
(comparative "synthetic (-er / -st, with umlaut)")
;; ── negation ───────────────────────────────────────────────────────────
(negation-markers (nicht kein)) ; kein- negates an indefinite NP; nicht else
(negation-faithful yes) ; SACRED: polarity never dropped/inverted -> FLAG
;; ── lexicon provenance ─────────────────────────────────────────────────
(lexicon-source "UniMorph deu (primary) + kaikki.org German (gender override)")
(lexicon-license "CC-BY-SA 3.0 / GFDL"))
-41
View File
@@ -1,41 +0,0 @@
;;; lang_profile_en.el — English language profile for ELP.
;;; Mirrors lang_profile_es / lang_profile_pt; keys the realizer's construction
;;; switches. English is typologically distinct from the Romance builds, so the
;;; flags differ where the grammar differs.
(lang_profile_en
(language "English")
(iso639 "en")
(family "Germanic")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop no) ; OBLIGATORY subjects — missing subject is FLAGGED
(obligatory-subject yes)
(grammatical-gender no) ; natural gender only (he/she/it), no NP agreement
(do-support yes) ; negation & questions of lexical verbs insert do/does/did
(subject-aux-inversion yes) ; yes/no + non-subject wh questions invert the operator
(article-selection "a/an/the") ; a/an resolved PHONOLOGICALLY (an hour, a university)
(adjective-position prenominal) ; attributive adjectives precede the noun; invariant
(has-tag-questions yes) ; "...doesn't he?" — operator + reversed polarity
(has-there-existential yes) ; "there is/are/have been ..."
(possessive-clitic "'s") ; saxon genitive; plural in -s -> bare apostrophe
(question-punct plain) ; ? and ! only (no inverted marks)
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "3sg-present-only") ; only 3sg present -s (+ suppletive be)
(auxiliary-order (modal perfect progressive passive main))
(perfect-aux "have") ; have + past participle
(progressive-aux "be") ; be + present participle
(passive-aux "be") ; be + past participle (+ by-agent)
(future "will + base") ; no inflectional future
(comparative "synthetic-or-periphrastic") ; -er/-est vs more/most by syllables
;; ── SACRED safety bar (shared with es/pt) ──────────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
;; ── DIALECT overlay (post-realization, one core -> US/UK/AU) ────────────
(dialect US) ; default; profile field switches the overlay
(dialects (US UK AU))
(dialect-canonical US) ; core is authored in US orthography
(dialect-overlay "dialect_en.to_dialect") ; orthography + lexis + grammar prefs
(dialect-covers (spelling lexis collective-agreement gotten/got)))
-45
View File
@@ -1,45 +0,0 @@
;;; lang_profile_es.el — Spanish language profile for ELP.
;;; Keys the realizer's construction switches. Mirrors lang_profile_en / _pt.
(lang_profile_es
(language "Spanish")
(iso639 "es")
(family "Romance")
;; -- core typology flags -------------------------------------------------
(pro-drop yes) ; subjects routinely dropped; agreement carries person
(obligatory-subject no)
(grammatical-gender yes) ; m/f on every noun; article+adjective AGREE
(gender-source lexicon); REAL per-noun gender from UniMorph — NOT a heuristic
(do-support no)
(subject-aux-inversion no) ; questions by intonation/punctuation, not inversion
(question-strategy intonation)
(article-selection "el/la/los/las un/una/unos/unas")
(stressed-a-rule yes) ; fem sg noun in stressed a-/ha- takes el/un (el agua)
(adjective-position postnominal) ; default post; a few prenominal + apocope
(adjective-agreement "gender+number")
(question-punct inverted) ; opening ¿ ¡ required
;; -- MANDATORY CONTRACTIONS (coordinator quality bar) --------------------
(contractions ((de el "del") (a el "al")))
(contraction-mandatory yes) ; 'de el'/'a el' MUST surface as del/al
;; -- verb / aspect system ------------------------------------------------
(verb-classes (ar er ir))
(tenses (present preterite imperfect future conditional))
(moods (ind sbjv imp))
(finite-agreement "person+number (6 slots)")
(perfect-aux "haber") ; haber + past participle (invariant -o)
(progressive-aux "estar") ; estar + gerund
(passive-aux "ser") ; ser + participle (agrees) + por-agent
(copula-split "ser/estar") ; permanent vs stage-level
(future "infinitive + é/ás/á/emos/éis/án")
;; -- clitics / government ------------------------------------------------
(object-clitics yes) ; me te lo la le nos os los las; proclisis/enclisis
(clitic-order "se II I III (le+lo -> se lo)")
(enclisis "imperative/infinitive/gerund + accent repair (dá+me+lo->dámelo)")
(verb-prep-government yes) ; verbs select prep (protestar+contra, escapar+de)
;; -- SACRED safety bar (shared with en/pt) -------------------------------
(negation-faithful yes)) ; polarity never dropped/inverted; unplaceable -> FLAG
-74
View File
@@ -1,74 +0,0 @@
;;; lang_profile_fr.el — French language profile for ELP.
;;; Mirrors lang_profile_it / lang_profile_es; keys the realizer's construction
;;; switches. French is a Romance sibling (~54% of the realizer code and the whole
;;; clause-engine architecture reused), but carries the family's biggest surface
;;; deltas: NOT pro-drop, DISCONTINUOUS negation, and an orthography/phonology
;;; mismatch (elision, liaison) that makes exact-match genuinely hard.
(lang_profile_fr
(language "French")
(iso639 "fr")
(family "Romance")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop no) ; << French-specific: subject clitic OBLIGATORY
(obligatory-subject yes) ; je/tu/il/elle/nous/vous/ils/elles always overt
(grammatical-gender yes) ; m/f; full NP agreement (art + adj + participle)
(do-support no)
(subject-aux-inversion optional) ; est-ce que (default) OR clitic inversion (vas-tu)
(article-selection "le/la/l'/les ; un/une/des ; PARTITIVE du/de la/de l'/des")
(article-drives-contraction yes) ; à+le=au, de+le=du feed off article choice
(adjective-position "postnominal-default + prenominal-BAGS") ; beau/bon/grand/
; petit/jeune/vieux/nouveau + ordinals prenominal
; (beau->bel, nouveau->nouvel, vieux->vieil / vowel)
(question-punct "space-before") ; French typography: ' ?' ' !' (no ¿¡)
;; ── elision (orthography/phonology mismatch — French-specific) ──────────
(elision ((le l') (la l') (je j') (ne n') (de d') (que qu')
(me m') (te t') (se s') (ce c'))) ; before vowel / h-muet
(elision-h-muet yes) ; l'homme, l'hôpital (h-aspiré exception list kept)
(liaison noted-not-modeled) ; phonological, not written in surface
;; ── MANDATORY prep+article contractions ────────────────────────────────
(contractions ((à le au) (à les aux) (de le du) (de les des)))
(contraction-mandatory yes) ; *à le -> au obligatory; à la / à l' uncontracted
(partitive ((m-sg du) (f-sg "de la") (vowel "de l'") (pl des)))
(partitive-under-neg "de") ; << gap in current build: 'ne … pas de pain'
;; ── clitic system ──────────────────────────────────────────────────────
(clitics yes)
(clitic-order (me te se nous vous | le la les | lui leur | y | en))
(clitic-placement ((finite proclitic) ; je le lui donne
(imperative-affirmative enclitic-hyphen) ; donne-le-moi
(imperative-negative "ne+proclitic+verb+pas") ; ne le donne pas
(infinitive enclitic))) ; PARTIAL: clitic-climbing
; onto infinitive under modal
(clitic-imperative-shift ((me moi) (te toi))) ; final me/te -> moi/toi (donne-moi)
(clitic-particles (y en)) ; locative y, partitive/genitive en
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "person+number (written; many homophones)")
(tenses (présent imparfait passé-simple futur conditionnel
subjonctif-présent subjonctif-imparfait impératif))
(compound-past "passé-composé = aux(present) + participe passé")
(perfect-aux "être/avoir (LEXICAL selection)") ; << French-specific
(etre-aux-class "intransitive motion/change (aller venir arriver partir
entrer sortir monter descendre naître mourir rester
tomber retourner passer devenir revenir rentrer) + ALL
pronominal verbs")
(participle-agreement ((être subject) ; elle est allée / elles venues
(avoir preceding-direct-object))) ; je les ai vus
(progressive "être en train de + infinitif") ; no dedicated aux
(copula "être (single; no ser/estar, no essere/stare)")
(passive-aux "être (+ par-agent)")
(future inflectional) ; parlera, sera
(comparative "plus/moins ADJ que")
(superlative "le/la plus ADJ (de …)") ; PARTIAL word-order in build
;; ── SACRED safety bar (shared with es/pt/it/en) ────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
(negation "DISCONTINUOUS: ne (preverbal) … pas/jamais/rien/personne/
plus/guère/que (postverbal)") ; << biggest structural delta
(negation-ne-elides yes) ; ne -> n' before vowel (n'ai pas vu)
(negation-passe-composé "ne + aux + pas + participe") ; n'ai pas vu
(negative-concord partial)) ; personne/rien as arguments post-participle
-70
View File
@@ -1,70 +0,0 @@
;;; lang_profile_it.el — Italian language profile for ELP.
;;; Mirrors lang_profile_es / lang_profile_pt; keys the realizer's construction
;;; switches. Italian is a Romance sibling, so ~85% of the flags match ES/PT; the
;;; essere/avere auxiliary split and phonological article selection are the deltas.
(lang_profile_it
(language "Italian")
(iso639 "it")
(family "Romance")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop yes) ; null subjects default; overt pronoun = emphatic
(obligatory-subject no)
(grammatical-gender yes) ; m/f; full NP agreement (art + adj + participle)
(do-support no)
(subject-aux-inversion no) ; yes/no Q = declarative order + '?'; no inversion
(article-selection "il/lo/l'/i/gli + la/l'/le ; un/uno/un'/una") ; PHONOLOGICAL:
; lo/gli/uno before s+cons, z, gn, ps, pn, x, y, i+V;
; l'/un' before a vowel (elision, glued to next word)
(article-drives-contraction yes) ; article choice feeds the prep+art contraction
(adjective-position "postnominal-default + prenominal-class") ; bello/buono/grande
; /nuovo/vecchio/primo... prenominal (with apocope)
(question-punct plain) ; ? and ! only (no inverted ¿ ¡)
;; ── MANDATORY prep+article contractions ────────────────────────────────
(contractions ((di il del) (di lo dello) (di la della) (di i dei)
(di gli degli) (di le delle) (di l' dell')
(a il al) (a lo allo) (a la alla) (a i ai) (a gli agli)
(a le alle) (a l' all')
(da il dal) (da la dalla) (da gli dagli) (da l' dall')
(in il nel) (in la nella) (in gli negli) (in l' nell')
(su il sul) (su la sulla) (su gli sugli) (su l' sull')))
(contraction-mandatory yes) ; *di il -> del is obligatory, never uncontracted
(prep-no-contract (per tra fra)) ; per la strada (NOT *perla)
;; ── clitic system ──────────────────────────────────────────────────────
(clitics yes)
(clitic-placement ((finite proclitic) ; lo vedo, non me lo dà
(imperative-affirmative enclitic) ; dammelo, guardalo
(imperative-negative-tu non+infinitive) ; non parlare / non lo fare
(infinitive enclitic) ; vederlo, aiutarmi (drop -e)
(gerund enclitic))) ; dandolo
(clitic-combination ((mi lo "me lo") (ti lo "te lo") (ci lo "ce lo")
(vi lo "ve lo") (si lo "se lo")
(gli lo "glielo") (le lo "glielo"))) ; glielo = ONE word
(clitic-particles (ci ne)) ; locative ci, partitive ne
(raddoppiamento (da fa di va sta)) ; monosyllabic imper double clitic: dammelo
;; ── verb / aspect system ───────────────────────────────────────────────
(finite-agreement "person+number (6-way)")
(tenses (presente imperfetto passato-remoto futuro condizionale
congiuntivo-presente congiuntivo-imperfetto imperativo))
(compound-past "passato-prossimo = aux(present) + participle")
(perfect-aux "essere/avere (LEXICAL selection)") ; << Italian-specific
(essere-aux-class unaccusative) ; motion/change-of-state/copular/pronominal
; (andare venire nascere morire diventare piacere
; + ALL reflexives) -> essere
(participle-agreement ((essere subject) ; è andata / sono arrivati
(avere preceding-acc-clitic))) ; li ho visti
(progressive-aux "stare + gerundio") ; sto parlando
(copula "essere (default) / stare (state: sto bene)")
(passive-aux "essere / venire (+ da-agent)")
(future inflectional) ; parlerò, sarà
(comparative "più/meno ADJ di")
;; ── SACRED safety bar (shared with es/pt/en) ───────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
(negation "non (preverbal) + concord") ; non...niente/nessuno/mai/più
(negative-concord yes) ; preverbal negative word (nessuno/niente) suppresses non
(neg-adverb-position between-aux-and-participle)) ; non ho MAI visto
-30
View File
@@ -1,30 +0,0 @@
;;; lang_profile_la.el — Latin language profile for ELP.
;;; Keys the realizer's construction switches. Companion to morphology-la.el.
(lang_profile_la
(language "Latin")
(iso639 "la")
(family "Italic")
;; -- core typology flags -------------------------------------------------
(pro-drop yes) ; person carried by verb ending; subjects dropped
(obligatory-subject no)
(grammatical-gender yes) ; m/f/n; adjective AGREES in case+gender+number
(gender-source lexicon) ; REAL per-noun gender from UniMorph lat
(articles none) ; Latin has no articles
(case-system yes) ; NOM GEN DAT ACC ABL VOC (+ rare LOC)
(cases (nom gen dat acc abl voc))
(word-order "SOV (default; free order, case-marked)")
(adjective-position "either (case agreement carries the link)")
(adjective-agreement "case+gender+number")
;; -- verb / aspect system ------------------------------------------------
(verb-classes (1 2 3 3io 4)) ; four conjugations + i-stem 3rd
(tenses (present imperfect future perfect pluperfect futureperfect))
(moods (indicative subjunctive imperative infinitive))
(voices (active passive))
(finite-agreement "person+number (6 slots)")
(citation "principal parts: pres-1sg / pres-inf / perf-participle")
;; -- SACRED safety bar ---------------------------------------------------
(negation-faithful yes)) ; polarity never dropped/inverted
-40
View File
@@ -1,40 +0,0 @@
;;; lang_profile_pt.el — Portuguese language profile for ELP.
;;; Keys the realizer's construction switches. Mirrors lang_profile_es.
(lang_profile_pt
(language "Portuguese")
(iso639 "pt")
(family "Romance")
;; -- core typology flags -------------------------------------------------
(pro-drop yes) ; subjects routinely dropped; agreement carries person
(obligatory-subject no)
(grammatical-gender yes) ; m/f on every noun; article+adjective AGREE
(gender-source lexicon) ; REAL per-noun gender from UniMorph por / kaikki
(do-support no)
(subject-aux-inversion no)
(question-strategy intonation)
(article-selection "o/a/os/as um/uma/uns/umas")
(adjective-position postnominal)
(adjective-agreement "gender+number")
;; -- MANDATORY CONTRACTIONS (prep + article) -----------------------------
(contractions ((de o "do") (de a "da") (em o "no") (em a "na")
(a o "ao") (a a "à") (por o "pelo") (por a "pela")))
(contraction-mandatory yes)
;; -- verb / aspect system ------------------------------------------------
(verb-classes (ar er ir))
(tenses (present preterite imperfect future conditional))
(moods (ind sbjv imp))
(finite-agreement "person+number (6 slots)")
(perfect-aux "ter") ; ter + past participle
(copula-split "ser/estar")
(personal-infinitive yes) ; distinctive PT inflected infinitive
;; -- clitics / government ------------------------------------------------
(object-clitics yes) ; mesoclisis/enclisis/proclisis by context
(verb-prep-government yes)
;; -- SACRED safety bar ---------------------------------------------------
(negation-faithful yes))
-71
View File
@@ -1,71 +0,0 @@
;;; lang_profile_ro.el — Romanian language profile for ELP.
;;; Romanian is the BIG typological delta of the Romance family. The verb/clause
;;; engine and the SACRED negation contract mirror the ES/PT/IT core, but the
;;; NOMINAL system is genuinely new: a SUFFIXED definite article, preserved CASE,
;;; a NEUTER gender, and a VOCATIVE. Those flags mark where the shared engine was
;;; extended rather than reused.
(lang_profile_ro
(language "Romanian")
(iso639 "ro")
(family "Romance (Eastern / Balkan)")
;; ── core typology flags ────────────────────────────────────────────────
(pro-drop yes) ; null subjects default; overt pronoun = emphatic
(obligatory-subject no)
(grammatical-gender yes) ; m / f / NEUTER (n)
(neuter-gender yes) ; << ROMANIAN-SPECIFIC: masc-agreeing SG, fem-agreeing PL
; (un tren nou / două trenuri noi)
(do-support no)
(subject-aux-inversion no) ; yes/no Q = declarative order + '?'
(question-punct plain) ; ? and ! only
;; ── SUFFIXED DEFINITE ARTICLE (the headline engine extension) ───────────
(definite-article suffixed) ; << UNIQUE IN ROMANCE: enclitic on the noun
(definite-forms ((m/n sg "-ul / -le / -l : om->omul, câine->câinele, codru->codrul")
(f sg "-a / -ea / -ua : casă->casa, carte->cartea, stea->steaua")
(m pl "-i : oameni->oamenii")
(f/n pl "-le : case->casele, trenuri->trenurile")))
(article-host ((no-prenom-adj noun) ; omul bun
(prenom-adj adjective))) ; bunul om (adj carries the article)
(indefinite-article ((m/n "un") (f "o") (pl "niște") (gen/dat-pl "unor")))
;; ── CASE (preserved; NOM/ACC vs GEN/DAT) ────────────────────────────────
(case (nom/acc gen/dat vocative)) ; << ROMANIAN-SPECIFIC
(case-syncretism "nom=acc ; gen=dat")
(genitive-marking "gen/dat definite: -lui (m/n), -ei/-i (f), -lor (pl)")
(genitival-article ((m sg "al") (f sg "a") (m pl "ai") (f/n pl "ale"))) ; o carte a lui
(possession "definite-head + gen/dat possessor: casa băiatului")
(vocative ((m sg "-ule/-e : omule, băiete") (f sg "-o : Mario, fato")
(pl "-lor")))
;; ── verb / aspect system ────────────────────────────────────────────────
(finite-agreement "person+number (6-way)")
(tenses (prezent imperfect perfect-simplu conjunctiv-prezent
imperativ (periphrastic: perfect-compus viitor conditional)))
(compound-past "perfectul compus = a-avea-clitic + INVARIABLE participle")
(perfect-aux "a avea (am/ai/a/am/ați/au) — ONE auxiliary for ALL verbs")
(perfect-aux-split no) ; << SIMPLER than Italian: no essere/avere selection
(participle-agreement none) ; invariable in the perfect compus (agrees only as
; an adjective / in the passive)
(future "voi/vei/va/vom/veți/vor + infinitive (viitor literar)")
(conditional "aș/ai/ar/am/ați/ar + infinitive")
(subjunctive "conjunctiv: particle 'să' + subjunctive present")
(modal-complement "modal + să + subjunctive (vreau să merg, poți să ajuți)")
(copula "a fi")
(passive "a fi + participle (participle AGREES like an adjective)")
(comparative "mai / mai puțin ADJ decât")
;; ── clitic system (partial — see honest gaps) ───────────────────────────
(clitics yes)
(clitic-set ((acc te îl o ne îi le) (dat îmi îți îi ne le)
(refl te se ne se)))
(clitic-placement ((finite proclitic) ; îmi place, o văd
(perfect-compus elision) ; << m-am, l-am, i-am (PARTIAL)
(imperative-affirmative enclitic))) ; dă-mi (PARTIAL)
;; ── SACRED safety bar (shared with es/pt/it/en) ─────────────────────────
(negation-faithful yes) ; polarity never dropped/inverted; unplaceable -> FLAG
(negation "nu (single preverbal marker) + concord")
(negative-concord yes) ; nu … nimic / nimeni / niciodată / niciun
(negative-imperative "nu + INFINITIVE : nu pleca! (KNOWN GAP: uses imperative stem)"))
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
-100
View File
@@ -1,100 +0,0 @@
# -*- coding: utf-8 -*-
"""Full-lexicon vocabulary-{de,la}.el emitters (custom field mapping for the
German declension/gender API and the Latin case-paradigm API). Reuses the
chunked seed-fn writer from gen_elp_seed_full.
"""
import sys, importlib
from gen_elp_seed_full import write_seed
def uw(x):
"""Unwrap (form, source) tuples that some morphology fns return."""
if isinstance(x, (tuple, list)):
return x[0] if x else ""
return x if x is not None else ""
def build_de():
M = importlib.import_module("morphology_de_full")
rows = []; st = {"verbs":0,"nouns":0,"adjs":0}
# nouns: form0=nom-sg(lemma) form1=plural form2=gender
for lem in sorted(M._NOUNS):
if not lem: continue
try:
g = uw(M.noun_gender(lem))
pl = uw(M.pluralize(lem))
except Exception:
continue
rows.append([lem, "noun", lem, pl, g or "", "", "gender:lexicon"])
st["nouns"] += 1
# adjs: form0=positive form1=comparative form2=superlative
for lem in sorted(M._ADJS):
if not lem: continue
try:
cmpr = uw(M.comparative(lem))
sprl = uw(M.superlative(lem))
except Exception:
continue
rows.append([lem, "adj", lem, cmpr, sprl, "", "degree:lexicon"])
st["adjs"] += 1
# verbs (only the ~30 irregular/strong stems the cache carries):
# form0=pres-3sg form1=past-3sg form2=past-participle
if hasattr(M, "_VERBS"):
for lem in sorted({k[0] if isinstance(k, tuple) else k for k in M._VERBS}):
if not lem: continue
try:
f0 = uw(M.finite(lem, "present", "third", "singular"))
f1 = uw(M.finite(lem, "past", "third", "singular"))
pp = uw(M.past_participle(lem))
except Exception:
continue
rows.append([lem, "verb", f0, f1, pp, "", "class:strong/irregular"])
st["verbs"] += 1
return rows, st
def build_la():
M = importlib.import_module("morphology_lat_full")
rows = []; st = {"verbs":0,"nouns":0,"adjs":0}
def dn(lem, c, n):
try:
r = M.decline_noun(lem, c, n)
return uw(r)
except Exception:
return ""
# nouns: dictionary citation — form0=nom-sg form1=gen-sg form2=gender
for lem in sorted(M._NOUNS):
if not lem: continue
nom = dn(lem, "NOM", "SG") or lem
gen = dn(lem, "GEN", "SG")
try: g = uw(M.noun_gender(lem))
except Exception: g = ""
rows.append([lem, "noun", nom, gen, g, "", "case-paradigm nom/gen-sg"])
st["nouns"] += 1
# adjs: three-gender nom-sg citation — form0=masc form1=fem form2=neut
for lem in sorted(M._ADJS):
if not lem: continue
try:
m = uw(M.decline_adj(lem, "NOM", "MASC", "SG")) or lem
f = uw(M.decline_adj(lem, "NOM", "FEM", "SG"))
nt = uw(M.decline_adj(lem, "NOM", "NEUT", "SG"))
except Exception:
continue
rows.append([lem, "adj", m, f, nt, "", "3-gender nom-sg"])
st["adjs"] += 1
# verbs: principal parts — form0=pres-ind-1sg form1=pres-infinitive form2=perf-participle
if hasattr(M, "_VERBS"):
for lem in sorted({k[0] if isinstance(k, tuple) else k for k in M._VERBS}):
if not lem: continue
try:
f0 = uw(M.conjugate(lem, "present", "indicative", "active", "first", "singular"))
inf = uw(M.infinitive(lem, "present", "active"))
pp = uw(M.participle(lem, "perfect", "nom", "m", "singular"))
except Exception:
continue
rows.append([lem, "verb", f0, inf, pp, "", "principal-parts pres1sg/inf/pfppl"])
st["verbs"] += 1
return rows, st
if __name__ == "__main__":
lang = sys.argv[1]; out = sys.argv[2]
rows, st = build_de() if lang == "de" else build_la()
total, _ = write_seed(lang, rows, st, out)
print(f"{lang}: wrote {out} total={total} verbs={st['verbs']} nouns={st['nouns']} adjs={st['adjs']}")
-129
View File
@@ -1,129 +0,0 @@
# -*- coding: utf-8 -*-
"""gen_elp_seed_full.py — emit a FULL-lexicon vocabulary-{lang}.el in the
established ELP seed-fn format (same as vocabulary-non.el / the 18 classical
languages), iterating the ENTIRE morphology_{lang}_full lexicon (every verb,
noun, adjective lemma) — NOT a curated demo core.
Schema per row: [lemma, pos, form0, form1, form2, en_translation, semantic_hint]
Verbs: form0=pres-ind-3sg form1=preterite-3sg form2=past-participle
Nouns: form0=singular form1=plural form2=REAL gender (lexicon)
Adjs : form0=masc-sg form1=fem-sg form2=masc-pl
Output structure (chunked to stay within the proven ~5k-append/function scale):
fn vocab_{lang}_seed_pN(v) -> [[String]] { ... appends ... return v }
fn vocab_{lang}_seed() -> [[String]] { chains all chunks; return v }
fn vocab_{lang}_lookup(w) -> [String] { linear scan }
Usage: python3 gen_elp_seed_full.py <lang> <out.el>
"""
import sys, importlib
CHUNK = 5000
def esc(s):
return str(s).replace("\\", "\\\\").replace('"', '\\"')
def row(fields):
return " let v = native_list_append(v, [" + ", ".join(f'"{esc(f)}"' for f in fields) + "])"
def build_rows(lang, M):
rows = []
stats = {"verbs":0,"nouns":0,"adjs":0}
has = lambda n: hasattr(M, n)
# --- verbs ---
if has("_VERBS") and has("conjugate"):
verbs = sorted({k[0] for k in M._VERBS})
for lem in verbs:
if not lem: continue
try:
f0, s0 = M.conjugate(lem, "ind", "present", "third", "singular")
f1, _ = M.conjugate(lem, "ind", "preterite", "third", "singular")
pp, _ = (M.participle(lem) if has("participle") else ("",""))
except Exception:
continue
vclass = lem[-2:] if lem[-2:] in ("ar","er","ir","re") else lem[-2:]
rows.append([lem, "verb", f0 or "", f1 or "", pp or "", "", "class:"+vclass+" src:"+str(s0)])
stats["verbs"] += 1
# --- nouns ---
if has("_NOUNS") and has("inflect_noun"):
for lem in sorted(M._NOUNS):
if not lem: continue
try:
sg, _ = M.inflect_noun(lem, "singular")
pl, _ = M.inflect_noun(lem, "plural")
g = M.noun_gender(lem) if has("noun_gender") else ""
except Exception:
continue
src = "lexicon" if (isinstance(M._NOUNS.get(lem), dict) and M._NOUNS[lem].get("g")) else "heuristic"
rows.append([lem, "noun", sg or lem, pl or "", g or "", "", "gender:"+src])
stats["nouns"] += 1
# --- adjectives ---
if has("_ADJS") and has("inflect_adj"):
for lem in sorted(M._ADJS):
if not lem: continue
try:
m_sg, _ = M.inflect_adj(lem, "m", "singular")
f_sg, _ = M.inflect_adj(lem, "f", "singular")
m_pl, _ = M.inflect_adj(lem, "m", "plural")
except Exception:
continue
rows.append([lem, "adj", m_sg or lem, f_sg or "", m_pl or "", "", "src:lexicon"])
stats["adjs"] += 1
return rows, stats
def write_seed(lang, rows, stats, out_path):
"""Write vocabulary-{lang}.el in the chunked seed-fn format from prebuilt rows.
Each row is a 7-field list [lemma,pos,f0,f1,f2,gloss,hint]."""
total = len(rows)
chunks = [rows[i:i+CHUNK] for i in range(0, total, CHUNK)] or [[]]
L = []
L.append(f"// vocabulary-{lang}.el — FULL {lang} lexicon for ELP surface realization.")
L.append(f"// Generated by gen_elp_seed_full.py from morphology_{lang}_full")
L.append(f"// (real UniMorph + kaikki.org Wiktionary forms; gender from lexicon, not heuristic).")
L.append(f"// Entries: {total} (verbs={stats['verbs']} nouns={stats['nouns']} adjs={stats['adjs']})")
L.append(f"// Schema: [lemma, pos, form0, form1, form2, en_translation, semantic_hint]")
L.append(f"// verbs: form0=pres-3sg form1=pret-3sg form2=past-participle")
L.append(f"// nouns: form0=sg form1=pl form2=REAL gender adjs: form0=m-sg form1=f-sg form2=m-pl")
L.append("")
for ci, ch in enumerate(chunks):
L.append(f"fn vocab_{lang}_seed_p{ci}(v: [[String]]) -> [[String]] {{")
for r in ch:
L.append(row(r))
L.append(" return v")
L.append("}")
L.append("")
L.append(f"fn vocab_{lang}_seed() -> [[String]] {{")
L.append(" let v: [[String]] = native_list_empty()")
for ci in range(len(chunks)):
L.append(f" let v = vocab_{lang}_seed_p{ci}(v)")
L.append(" return v")
L.append("}")
L.append("")
L.append(f"fn vocab_{lang}_lookup(word: String) -> [String] {{")
L.append(f" let vocab: [[String]] = vocab_{lang}_seed()")
L.append(" let n: Int = native_list_len(vocab)")
L.append(" let i: Int = 0")
L.append(" while i < n {")
L.append(" let entry: [String] = native_list_get(vocab, i)")
L.append(' if str_eq(native_list_get(entry, 0), word) { return entry }')
L.append(" let i = i + 1")
L.append(" }")
L.append(" return native_list_empty()")
L.append("}")
with open(out_path, "w", encoding="utf-8") as fh:
fh.write("\n".join(L) + "\n")
return total, stats
def emit(lang, out_path):
M = importlib.import_module(f"morphology_{lang}_full")
rows, stats = build_rows(lang, M)
return write_seed(lang, rows, stats, out_path)
if __name__ == "__main__":
lang, out = sys.argv[1], sys.argv[2]
total, stats = emit(lang, out)
print(f"{lang}: wrote {out} total={total} verbs={stats['verbs']} nouns={stats['nouns']} adjs={stats['adjs']}")
-572
View File
@@ -1,572 +0,0 @@
# -*- coding: utf-8 -*-
"""morphology_ca_full.py — production-grade Catalan morphological generator.
Same design as morphology_it_full.py (its Romance sibling); Catalan-specific data.
VERBS
UniMorph Catalan (github.com/unimorph/cat, CC-BY-SA 3.0)
7,535 verb lemmas × paradigm, CLEAN orthography:
present, imperfet (PST;IPFV), pretèrit simple (PST;PFV), futur,
condicional (COND), subjuntiu present (SBJV;PRS) / imperfet (SBJV;PST),
imperatiu (POS;IMP), infinitiu (NFIN), gerundi (V.CVB;PRS),
participi (V.PTCP;PST) — WITH full gender+number agreement forms
(cantat/cantada/cantats/cantades) stored directly.
ca_irreg_verbs.json — verbs UniMorph MISSES or under-populates
(anar, fer, plus core auxiliaries ser/haver/estar/tenir…), extracted from
kaikki.org Catalan by build_ca_irreg.py. Priority layer. Supplies anar,
whose present (vaig/vas/va/anem/aneu/van) is ALSO the PERIPHRASTIC-PRETERITE
auxiliary (vaig cantar = 'I sang') — a hallmark Catalan construction.
NOUNS + ADJECTIVES — kaikki.org Catalan (Wiktionary extract, CC-BY-SA 3.0)
noun lemmas WITH inherent gender + real plural (resolved PER LEMMA).
adjective lemmas with real feminine + plural forms.
Fallbacks degrade, never crash:
verbs : regular -ar/-er/-re/-ir rule generator (+ -car/-gar/-çar spelling).
nouns : gender heuristic + rule pluralization (-a→-es with ç/c/g/j/qu/gu
spelling changes; sibilant-final → -os; else -s). Ambiguous → FLAG.
adjs : -o? no (Catalan masc often consonant/-e); fem -a rule + plural rule.
Confidence flag per form: "lexicon" | "rule" | "fallback" (low → FLAG).
Public API (used by realizer_ca.py):
conjugate(lemma, mood, tense, person, number) -> (form, conf)
peri_pret_aux(person, number) -> form # anar-present, for vaig+INF
participle(lemma, gender, number) -> (form, conf)
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"
inflect_noun(lemma, number, gender=None) -> (form, conf)
inflect_adj(lemma, gender, number) -> (form, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "cat.unimorph")
_IRREG = os.path.join(_HERE, "data", "ca_irreg_verbs.json")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_ca.jsonl")
_CACHE = os.path.join(_HERE, "data", "ca_morph_cache.pkl")
_VERB_KEYMAP = {
("ind", "present"): {"IND", "PRS"},
("ind", "imperfect"): {"IND", "PST", "IPFV"},
("ind", "preterite"): {"IND", "PST", "PFV"},
("ind", "future"): {"IND", "FUT"},
("ind", "conditional"): {"COND"},
("sbjv", "present"): {"SBJV", "PRS"},
("sbjv", "imperfect"): {"SBJV", "PST"},
("imp", "affirmative"): {"POS", "IMP"},
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── verbs from UniMorph ──────────────────────────────────────────────────────────
def _build_verbs():
verbs = {}
part = {} # lemma -> {("m","SG"):form, ("f","SG"):..., ("m","PL"):..., ("f","PL"):...}
ger = {}
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f:
g = "f" if "FEM" in f else "m"
n = "PL" if "PL" in f else "SG"
part.setdefault(lemma, {})[(g, n)] = form
continue
if head == "V.CVB":
if "PRS" in f:
ger.setdefault(lemma, form)
continue
if head != "V":
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
for (mood, tense), req in _VERB_KEYMAP.items():
if not req <= f:
continue
if tense == "imperfect" and "PFV" in f:
continue
if tense == "preterite" and "IPFV" in f:
continue
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), form)
break
return verbs, part, ger
# ── kaikki nouns + adjectives ────────────────────────────────────────────────────
_EXCL_FORM_TAGS = {"alternative", "archaic", "obsolete", "dialectal", "regional",
"diminutive", "augmentative", "pejorative", "comparative",
"superlative", "misspelling", "rare", "informal", "literary",
"poetic", "error-unrecognized-form", "Balearic", "Valencian",
"dated", "nonstandard"}
def _kaikki_gender(arg):
if not arg:
return None
a = str(arg).lower()
if a.startswith("f"):
return "f"
if a.startswith("m"):
return "m"
return None
def _build_nouns_adjs():
nouns = {}
adjs = {}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
pos = d.get("pos")
word = d.get("word", "")
if not word or " " in word:
continue
forms = d.get("forms", []) or []
if pos == "noun":
ht = d.get("head_templates") or []
g = None
if ht:
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
if g is None:
tags = d.get("tags") or []
if "feminine" in tags:
g = "f"
elif "masculine" in tags:
g = "m"
pl = None
for x in forms:
t = set(x.get("tags") or [])
if "plural" in t and not (t & _EXCL_FORM_TAGS):
fm = x.get("form")
if fm and " " not in fm and fm not in ("#", "", "-"):
pl = fm
break
if word not in nouns:
nouns[word] = {"g": g, "SG": word, "PL": pl}
else:
cur = nouns[word]
if cur.get("g") is None and g:
cur["g"] = g
if not cur.get("PL") and pl:
cur["PL"] = pl
elif pos == "adj":
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word)
for x in forms:
t = set(x.get("tags") or [])
fm = x.get("form")
if not fm or " " in fm or (t & _EXCL_FORM_TAGS):
continue
if "feminine" in t and "plural" in t:
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
elif "masculine" in t and "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
elif "feminine" in t:
d0[("f", "SG")] = d0.get(("f", "SG")) or fm
elif "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
return nouns, adjs
def _build_cache():
verbs, part, ger = _build_verbs()
nouns, adjs = _build_nouns_adjs()
with open(_IRREG, encoding="utf-8") as fh:
irreg = json.load(fh)
data = {"verbs": verbs, "part": part, "ger": ger,
"nouns": nouns, "adjs": adjs, "irreg": irreg}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [_UNIMORPH, _KAIKKI, _IRREG]
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PART, _GER, _NOUNS, _ADJS, _IRREGV = (
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"],
_LEX["irreg"])
_PERI = _IRREGV.get("_peri_pret_aux", {})
# ── regular verb rule fallback ───────────────────────────────────────────────────
def _vclass(lemma):
if lemma.endswith("ar"):
return "ar"
if lemma.endswith("re"):
return "re"
if lemma.endswith("er"):
return "er"
if lemma.endswith("ir"):
return "ir"
return None
# endings [1sg,2sg,3sg,1pl,2pl,3pl] — central Catalan
_REG = {
("ind", "present", "ar"): ["o", "es", "a", "em", "eu", "en"],
("ind", "present", "re"): ["o", "s", "", "em", "eu", "en"],
("ind", "present", "er"): ["o", "s", "", "em", "eu", "en"],
("ind", "present", "ir"): ["o", "es", "", "im", "iu", "en"], # pure -ir (dormir)
("ind", "imperfect", "ar"): ["ava", "aves", "ava", "àvem", "àveu", "aven"],
("ind", "imperfect", "re"): ["ia", "ies", "ia", "íem", "íeu", "ien"],
("ind", "imperfect", "er"): ["ia", "ies", "ia", "íem", "íeu", "ien"],
("ind", "imperfect", "ir"): ["ia", "ies", "ia", "íem", "íeu", "ien"],
("ind", "preterite", "ar"): ["í", "ares", "à", "àrem", "àreu", "aren"],
("ind", "preterite", "re"): ["í", "eres", "é", "érem", "éreu", "eren"],
("ind", "preterite", "er"): ["í", "eres", "é", "érem", "éreu", "eren"],
("ind", "preterite", "ir"): ["í", "ires", "í", "írem", "íreu", "iren"],
("sbjv", "present", "ar"): ["i", "is", "i", "em", "eu", "in"],
("sbjv", "present", "re"): ["i", "is", "i", "em", "eu", "in"],
("sbjv", "present", "er"): ["i", "is", "i", "em", "eu", "in"],
("sbjv", "present", "ir"): ["i", "is", "i", "im", "iu", "in"],
("sbjv", "imperfect", "ar"): ["és", "essis", "és", "éssim", "éssiu", "essin"],
("sbjv", "imperfect", "re"): ["és", "essis", "és", "éssim", "éssiu", "essin"],
("sbjv", "imperfect", "er"): ["és", "essis", "és", "éssim", "éssiu", "essin"],
("sbjv", "imperfect", "ir"): ["ís", "issis", "ís", "íssim", "íssiu", "issin"],
("imp", "affirmative", "ar"): [None, "a", "i", "em", "eu", "in"],
("imp", "affirmative", "re"): [None, "", "i", "em", "eu", "in"],
("imp", "affirmative", "er"): [None, "", "i", "em", "eu", "in"],
("imp", "affirmative", "ir"): [None, "", "i", "im", "iu", "in"],
}
_FUT = ["é", "às", "à", "em", "eu", "an"]
_COND = ["ia", "ies", "ia", "íem", "íeu", "ien"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _apply_ar_spelling(stem, ending):
"""-car/-gar/-çar/-jar spelling before front (e/i) endings."""
front = ending[:1] in ("e", "i", "é", "í")
if not front:
# ç before back vowel stays; but -çar stem already ends ç
return stem + ending
if stem.endswith("c"):
return stem[:-1] + "qu" + ending
if stem.endswith("g"):
return stem[:-1] + "gu" + ending
if stem.endswith("ç"):
return stem[:-1] + "c" + ending
if stem.endswith("j"):
return stem[:-1] + "g" + ending
if stem.endswith("qu"):
return stem + ending
return stem + ending
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
body = lemma[:-2]
i = _slot_idx(person, number)
if mood == "ind" and tense in ("future", "conditional"):
# future/cond stem = infinitive (for -re verbs drop final -e)
stem = lemma[:-1] if vc == "re" else lemma
end = (_FUT if tense == "future" else _COND)[i]
return stem + end
table = _REG.get((mood, tense, vc))
if not table:
return None
end = table[i]
if end is None:
return None
if vc == "ar":
return _apply_ar_spelling(body, end)
# -re/-er/-ir: guard double vowel
if body and body[-1:] == end[:1] and end[:1] in "":
return body[:-1] + end
return body + end
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
lemma = lemma.strip().lower()
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{number and number[:2].upper()}"
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{_NUMBER.get(number,'?')}"
# UniMorph (cleanly accented) takes priority; the kaikki irregulars layer is a
# FALLBACK for verbs/slots UniMorph lacks (anar, fer, and rarer paradigm cells).
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
if form:
return form, "lexicon"
ir = _IRREGV.get(lemma)
if ir and key in ir:
return ir[key], "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r is not None:
return r, "rule"
return lemma, "fallback"
def peri_pret_aux(person, number):
"""anar-present auxiliary for the periphrastic preterite (vaig cantar)."""
return _PERI.get(f"{_PERSON.get(person,'3')}|{_NUMBER.get(number,'SG')}", "va")
# ── PUBLIC: participle + gerund ──────────────────────────────────────────────────
def participle(lemma, gender="m", number="singular"):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
ir = _IRREGV.get(lemma)
base = None
if ir and "part" in ir:
# prefer explicit irregular agreement form (part_mSG/part_fSG/...)
exact = ir.get("part_" + g + num)
if exact:
return exact, "lexicon"
base = ir["part"]
elif lemma in _PART:
table = _PART[lemma]
if (g, num) in table:
return table[(g, num)], "lexicon"
base = table.get(("m", "SG"))
if base is None:
vc = _vclass(lemma)
if vc == "ar":
base = lemma[:-2] + "at"
elif vc == "ir":
base = lemma[:-2] + "it"
elif vc in ("er", "re"):
base = lemma[:-2] + "ut"
else:
return lemma, "fallback"
conf = "rule"
else:
conf = "lexicon"
# agreement on -t/-ut/-at/-it participles: m.sg base, f.sg +a (-da? no: -ada),
# Catalan: cantat/cantada/cantats/cantades; -t → f -da, pl -ts/-des
if base.endswith("t"):
stem = base[:-1]
forms = {"m|SG": base, "f|SG": stem + "da",
"m|PL": base + "s", "f|PL": stem + "des"}
return forms[f"{g}|{num}"], conf
if base.endswith("s"): # after sibilant participle (rare): pres->presa
stem = base
forms = {"m|SG": base, "f|SG": base + "a",
"m|PL": base + "os", "f|PL": base + "es"}
return forms[f"{g}|{num}"], conf
return base, conf
def gerund(lemma):
lemma = lemma.strip().lower()
ir = _IRREGV.get(lemma)
if ir and "ger" in ir:
return ir["ger"], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
vc = _vclass(lemma)
if vc == "ar":
return lemma[:-2] + "ant", "rule"
if vc in ("er", "re"):
return lemma[:-2] + "ent", "rule"
if vc == "ir":
return lemma[:-2] + "int", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
_FEM_SUF = ("ció", "sió", "tat", "tud", "esa", "esa", "dat", "ança", "ència",
"ància", "tud", "ícia", "esa", "or") # note -or is mixed; kaikki wins
_MASC_SUF = ("atge", "ment", " isme", "or")
def _gender_heuristic(noun):
for suf in ("ció", "sió", "tat", "tud", "esa", "ança", "ència", "ància",
"ícia", "etat"):
if noun.endswith(suf):
return "f"
if noun.endswith("a") and not noun.endswith("ma"):
return "f"
return "m"
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g") in ("m", "f"):
return d["g"]
return _gender_heuristic(lemma)
def _rule_plural(noun, gender):
"""Deterministic Catalan pluralization. (form, ok); ok=False FLAGS ambiguity."""
if not noun:
return noun, True
# stressed final vowel with accent → +ns (mà→mans is irregular; but capità→capitans)
if noun[-1:] in ("à", "é", "í", "ó", "ú"):
return noun + "ns", True
if noun.endswith("ça"):
return noun[:-2] + "ces", True # plaça→places
if noun.endswith("ca"):
return noun[:-2] + "ques", True # branca→branques
if noun.endswith("ga"):
return noun[:-2] + "gues", True # amiga→amigues
if noun.endswith("ja"):
return noun[:-2] + "ges", True # pluja→pluges
if noun.endswith("qua"):
return noun[:-3] + "qües", True
if noun.endswith("gua"):
return noun[:-3] + "gües", True
if noun.endswith("a"):
return noun[:-1] + "es", True # casa→cases
# sibilant-final → -os
if noun.endswith(("s", "ç", "x", "ig")) or noun.endswith(("ix", "tx", "tj")):
if noun.endswith("ç"):
return noun[:-1] + "ços", True # braç→braços
return noun + "os", True # peix→peixos, gas→gasos
if noun[-1:] in ("e", "i", "o", "u"):
return noun + "s", True
# consonant-final
return noun + "s", True
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if number == "singular":
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
if d and d.get("PL"):
return d["PL"], "lexicon"
g = gender or noun_gender(lemma)
form, ok = _rule_plural(lemma, g)
return form, ("rule" if ok else "fallback")
# ── PUBLIC: adjective agreement ──────────────────────────────────────────────────
def _fem_of(adj):
"""Regular Catalan feminine: consonant/-o? Catalan masc usually consonant or -e.
default +a with spelling changes; -e→-a for some; but many are invariable."""
a = adj
if a.endswith("a"):
return a
if a.endswith("e"):
return a[:-1] + "a" # ample→? actually 'ample' invariable; kaikki wins
if a.endswith("u"):
return a + "a"
if a.endswith("c"):
return a[:-1] + "ca" # ric→rica
if a.endswith("t"):
return a + "a" # alt→alta
return a + "a"
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
d = _ADJS.get(lemma)
if d:
form = d.get((g, num))
if form:
return form, "lexicon"
sg = d.get((g, "SG")) or d.get(("m", "SG")) or lemma
if num == "PL":
pl, ok = _rule_plural(sg, g)
return pl, ("rule" if ok else "fallback")
return sg, "lexicon"
# rule fallback
base = lemma if g == "m" else _fem_of(lemma)
if num == "SG":
return base, "rule"
pl, ok = _rule_plural(base, g)
return pl, ("rule" if ok else "fallback")
def lexicon_stats():
return {
"verb_source": "UniMorph Catalan (github.com/unimorph/cat) + kaikki.org "
"irregulars (anar/fer/auxiliaries)",
"noun_adj_source": "kaikki.org Catalan (Wiktionary extract)",
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
"unimorph_verb_forms": len(_VERBS),
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
"irregular_verb_lemmas": len([k for k in _IRREGV if not k.startswith("_")]),
"participle_lemmas": len(_PART),
"gerund_lemmas": len(_GER),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("cantar", "ind", "present", "first", "singular", "canto"),
("cantar", "ind", "present", "third", "plural", "canten"),
("ser", "ind", "present", "third", "singular", "és"),
("haver", "ind", "present", "first", "singular", "he"),
("anar", "ind", "present", "first", "singular", "vaig"),
("fer", "ind", "present", "third", "singular", "fa"),
("perdre", "ind", "present", "first", "singular", "perdo"),
("dormir", "ind", "present", "third", "plural", "dormen"),
("cantar", "ind", "future", "first", "singular", "cantaré"),
("cantar", "ind", "preterite", "third", "singular", "cantà"),
("tenir", "sbjv", "present", "first", "singular", "tingui"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
ok += got == exp
print(f" {flag}{lemma:8} {mood}/{tense:11} {per[:3]}.{num[:2]} -> {got:10} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" peri-pret anar: 1sg=", peri_pret_aux("first", "singular"),
"3pl=", peri_pret_aux("third", "plural"))
print(" gender casa=", noun_gender("casa"), "home=", noun_gender("home"),
"cavall=", noun_gender("cavall"), "cançó=", noun_gender("cançó"))
print(" plural casa->", inflect_noun("casa", "plural"),
"| plaça->", inflect_noun("plaça", "plural"),
"| peix->", inflect_noun("peix", "plural"),
"| braç->", inflect_noun("braç", "plural"),
"| home->", inflect_noun("home", "plural"))
print(" adj: alt/f/sg->", inflect_adj("alt", "f", "singular"),
"| bonic/f/pl->", inflect_adj("bonic", "f", "plural"),
"| vermell/f/sg->", inflect_adj("vermell", "f", "singular"))
print(" part: cantar/f/sg->", participle("cantar", "f", "singular"),
"| veure/f/pl->", participle("veure", "f", "plural"),
"| fer/m/sg->", participle("fer", "m", "singular"))
print(" ger: fer->", gerund("fer"), "| cantar->", gerund("cantar"))
-423
View File
@@ -1,423 +0,0 @@
# -*- coding: utf-8 -*-
"""morphology_de_full.py — production German morphological generator.
Real data, no toy tables:
PRIMARY — UniMorph German (github.com/unimorph/deu, CC-BY-SA 3.0).
~219k noun forms, ~199k verb forms. Supplies:
nouns : gender (MASC/FEM/NEUT) + case×number paradigm
(N;NOM/ACC/DAT/GEN; MASC/FEM/NEUT; SG/PL) — the genitive -(e)s,
dative-plural -n and the five plural classes are REAL forms, not
guessed.
verbs : full finite paradigm IND;{SG,PL};{1,2,3};{PRS,PST}, the past
participle (V.PTCP;PST, incl. reattached separable prefix
'zugefügt'), and — crucially for V2 — the SEPARATED finite form
UniMorph records directly ('füge zu', 'steht auf').
adjs : comparative / superlative (ADJ;CMPR, ADJ;SPRL).
SECONDARY — kaikki.org German (Wiktionary, CC-BY-SA/GFDL). Gap-fills noun
gender + plural where UniMorph is thin. Never overrides UniMorph.
Rule fallbacks (flagged 'rule'/'fallback') for lemmas absent from both lexicons:
present : -e/-st/-t/-en/-t/-en with e-epenthesis after -t/-d/-chn stems
plural : gender heuristic (fem -> -(e)n, else -e / umlaut left to lexicon)
ppart : weak ge-…-t
Adjective ENDINGS are rule-computed by the realizer (regular closed table);
this module only supplies the comparative/superlative STEM.
Perfect auxiliary (haben vs sein): sein for a curated set of intransitive
motion / change-of-state verbs (real German lexical property), else haben.
Public API:
noun_gender(lemma) -> 'm'|'f'|'n'
decline_noun(lemma, case, number) -> (form, conf)
pluralize(lemma) -> (form, conf)
finite(lemma, tense, person, number) -> (form, conf) # may contain ' prefix'
nonfinite(lemma, req) -> (form, conf) # req: 'inf'|'ppart'
past_participle(lemma) -> (form, conf)
separable_prefix(lemma) -> str|None
perfect_aux(lemma) -> 'haben'|'sein'
comparative(lemma)/superlative(lemma) -> (stem, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "deu.unimorph")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_de.jsonl")
_CACHE = os.path.join(_HERE, "data", "de_morph_cache.pkl")
_GENDER = {"MASC": "m", "FEM": "f", "NEUT": "n"}
# intransitive motion / change-of-state verbs that take SEIN in the perfect
_SEIN = {"gehen", "kommen", "fahren", "laufen", "rennen", "reisen", "fallen",
"steigen", "sinken", "wachsen", "sterben", "geschehen", "passieren",
"werden", "bleiben", "sein", "aufstehen", "einschlafen", "aufwachen",
"ankommen", "abfahren", "aufsteigen", "erscheinen", "verschwinden",
"fliegen", "schwimmen", "springen", "begegnen", "folgen", "gelingen",
"wandern", "ziehen", "flüchten", "eintreten", "einsteigen", "aussteigen"}
# hardcoded high-frequency irregular / auxiliary / modal paradigms (closed class,
# verified) — consulted before the lexicon so aux+modal chains are always correct.
_CORE = {
"sein": {"prs": {("first", "singular"): "bin", ("second", "singular"): "bist",
("third", "singular"): "ist", ("first", "plural"): "sind",
("second", "plural"): "seid", ("third", "plural"): "sind"},
"pst": {("first", "singular"): "war", ("second", "singular"): "warst",
("third", "singular"): "war", ("first", "plural"): "waren",
("second", "plural"): "wart", ("third", "plural"): "waren"},
"ppart": "gewesen"},
"haben": {"prs": {("first", "singular"): "habe", ("second", "singular"): "hast",
("third", "singular"): "hat", ("first", "plural"): "haben",
("second", "plural"): "habt", ("third", "plural"): "haben"},
"pst": {("first", "singular"): "hatte", ("second", "singular"): "hattest",
("third", "singular"): "hatte", ("first", "plural"): "hatten",
("second", "plural"): "hattet", ("third", "plural"): "hatten"},
"ppart": "gehabt"},
"werden": {"prs": {("first", "singular"): "werde", ("second", "singular"): "wirst",
("third", "singular"): "wird", ("first", "plural"): "werden",
("second", "plural"): "werdet", ("third", "plural"): "werden"},
"pst": {("first", "singular"): "wurde", ("second", "singular"): "wurdest",
("third", "singular"): "wurde", ("first", "plural"): "wurden",
("second", "plural"): "wurdet", ("third", "plural"): "wurden"},
"ppart": "geworden"},
}
_MODAL_PRS = {
"können": ("kann", "kannst", "kann", "können", "könnt", "können"),
"müssen": ("muss", "musst", "muss", "müssen", "müsst", "müssen"),
"wollen": ("will", "willst", "will", "wollen", "wollt", "wollen"),
"sollen": ("soll", "sollst", "soll", "sollen", "sollt", "sollen"),
"dürfen": ("darf", "darfst", "darf", "dürfen", "dürft", "dürfen"),
"mögen": ("mag", "magst", "mag", "mögen", "mögt", "mögen"),
}
_MODAL_PST = {
"können": ("konnte", "konntest", "konnte", "konnten", "konntet", "konnten"),
"müssen": ("musste", "musstest", "musste", "mussten", "musstet", "mussten"),
"wollen": ("wollte", "wolltest", "wollte", "wollten", "wolltet", "wollten"),
"sollen": ("sollte", "solltest", "sollte", "sollten", "solltet", "sollten"),
"dürfen": ("durfte", "durftest", "durfte", "durften", "durftet", "durften"),
"mögen": ("mochte", "mochtest", "mochte", "mochten", "mochtet", "mochten"),
}
_PN_ORDER = [("first", "singular"), ("second", "singular"), ("third", "singular"),
("first", "plural"), ("second", "plural"), ("third", "plural")]
_MODAL_PPART = {"können": "gekonnt", "müssen": "gemusst", "wollen": "gewollt",
"sollen": "gesollt", "dürfen": "gedurft", "mögen": "gemocht"}
for _m, _forms in _MODAL_PRS.items():
_CORE[_m] = {"prs": dict(zip(_PN_ORDER, _forms)),
"pst": dict(zip(_PN_ORDER, _MODAL_PST[_m])),
"ppart": _MODAL_PPART[_m]}
def _person_num(tags):
p = n = None
for t in tags:
if t in ("1", "2", "3"):
p = {"1": "first", "2": "second", "3": "third"}[t]
elif t == "SG":
n = "singular"
elif t == "PL":
n = "plural"
return p, n
def _build_from_unimorph():
nouns, verbs, adjs = {}, {}, {}
if not os.path.exists(_UNIMORPH):
return nouns, verbs, adjs
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tagstr = parts
tags = tagstr.split(";")
head = tags[0]
tset = set(tags)
if head == "N":
rec = nouns.setdefault(lemma, {"g": None, "cases": {}, "pl": None})
g = next((_GENDER[t] for t in tags if t in _GENDER), None)
if g and not rec["g"]:
rec["g"] = g
case = next((t for t in tags if t in ("NOM", "ACC", "DAT", "GEN")), None)
num = "plural" if "PL" in tset else ("singular" if "SG" in tset else None)
if case and num:
rec["cases"].setdefault((case, num), form)
if case == "NOM" and num == "plural" and not rec["pl"]:
rec["pl"] = form
elif head.startswith("V"):
rec = verbs.setdefault(lemma, {"prs": {}, "pst": {}, "ppart": None})
if "PTCP" in head and "PST" in tset:
rec["ppart"] = rec["ppart"] or form
elif "IND" in tset and ("PRS" in tset or "PST" in tset):
p, n = _person_num(tags)
if p and n:
slot = "prs" if "PRS" in tset else "pst"
rec[slot].setdefault((p, n), form)
elif head == "ADJ":
rec = adjs.setdefault(lemma, {})
if "CMPR" in tset:
rec.setdefault("cmpr", form.replace("am ", "").strip())
elif "SPRL" in tset:
rec.setdefault("sprl", form.replace("am ", "").replace("sten", "st")
if form.endswith("sten") else form.replace("am ", ""))
return nouns, verbs, adjs
def _build_from_kaikki(nouns):
"""Gap-fill noun gender + plural from kaikki German."""
if not os.path.exists(_KAIKKI):
return
_g = {"masculine": "m", "feminine": "f", "neuter": "n", "m": "m", "f": "f", "n": "n"}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
if d.get("pos") != "noun":
continue
w = d.get("word", "")
if not w or not w[0].isalpha() or " " in w:
continue
rec = nouns.setdefault(w, {"g": None, "cases": {}, "pl": None})
# GENDER: Wiktionary gender is hand-curated and OVERRIDES UniMorph's
# auto-tagged gender, which has known errors (e.g. UniMorph deu mis-
# records Zeit=MASC, Wagen=NEUT; Wiktionary has f, m correctly).
for h in d.get("head_templates", []) or []:
a = h.get("args", {}) or {}
raw = a.get("1") or a.get("g") or ""
code = str(raw).split(",")[0].strip().lower()
if code in _g:
rec["g"] = _g[code]
break
if not rec["pl"]:
for f in d.get("forms", []) or []:
t = set(f.get("tags", []) or [])
if "plural" in t and f.get("form") and "genitive" not in t:
rec["pl"] = f["form"]
break
def _build_cache():
nouns, verbs, adjs = _build_from_unimorph()
_build_from_kaikki(nouns)
data = {"nouns": nouns, "verbs": verbs, "adjs": adjs}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [p for p in (_UNIMORPH, _KAIKKI) if os.path.exists(p)]
newest = max((os.path.getmtime(p) for p in srcs), default=0)
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_NOUNS, _VERBS, _ADJS = _LEX["nouns"], _LEX["verbs"], _LEX["adjs"]
# ── nouns ────────────────────────────────────────────────────────────────────────
def noun_gender(lemma):
rec = _NOUNS.get(lemma) or _NOUNS.get(lemma.capitalize())
if rec and rec.get("g"):
return rec["g"]
# last-resort rule: -ung/-heit/-keit/-schaft/-tät/-ion -> f ; -chen/-lein -> n
low = lemma.lower()
if low.endswith(("ung", "heit", "keit", "schaft", "tät", "ion", "ik", "ei")):
return "f"
if low.endswith(("chen", "lein", "ment", "um")):
return "n"
return "m"
def pluralize(lemma):
rec = _NOUNS.get(lemma) or _NOUNS.get(lemma.capitalize())
if rec and rec.get("pl"):
return rec["pl"], "lexicon"
g = noun_gender(lemma)
if g == "f":
return (lemma + "en" if not lemma.endswith("e") else lemma + "n"), "rule"
return (lemma if lemma.endswith(("er", "en", "el")) else lemma + "e"), "rule"
def decline_noun(lemma, case, number):
"""case in NOM/ACC/DAT/GEN, number in singular/plural."""
rec = _NOUNS.get(lemma) or _NOUNS.get(lemma.capitalize())
if case == "DAT" and number == "singular":
# modern German drops the archaic dative -e ('dem Kinde' -> 'dem Kind');
# the article carries the case. Keep bare nominative form.
base = (rec or {}).get("cases", {}).get(("NOM", "singular")) or lemma
return base, ("lexicon" if rec else "rule")
if rec and rec.get("cases", {}).get((case, number)):
return rec["cases"][(case, number)], "lexicon"
if number == "plural":
pl, c = pluralize(lemma)
if case == "DAT" and not pl.endswith("n") and not pl.endswith("s"):
return pl + "n", c # dative plural -n
return pl, c
# singular
g = noun_gender(lemma)
if case == "GEN" and g in ("m", "n"):
return (lemma + "es" if lemma.endswith(("s", "ß", "z", "x")) else lemma + "s"), "rule"
return lemma, "lexicon" if rec else "rule"
# ── verbs ──────────────────────────────────────────────────────────────────────--
_PRS_ENDINGS = {("first", "singular"): "e", ("second", "singular"): "st",
("third", "singular"): "t", ("first", "plural"): "en",
("second", "plural"): "t", ("third", "plural"): "en"}
def _stem(lemma):
if lemma.endswith("en"):
return lemma[:-2]
if lemma.endswith("n"):
return lemma[:-1]
return lemma
def separable_prefix(lemma):
"""Return the separable prefix if the lemma is a separable-prefix verb."""
rec = _VERBS.get(lemma)
if rec:
for (_p, _n), form in rec.get("prs", {}).items():
if " " in form:
return form.rsplit(" ", 1)[1]
_SEP = ("auf", "aus", "ab", "an", "ein", "mit", "nach", "vor", "zu", "zurück",
"weg", "hin", "her", "los", "bei", "fest", "fort", "um", "zusammen")
_INSEP = ("be", "ge", "er", "ver", "zer", "ent", "emp", "miss")
for p in sorted(_SEP, key=len, reverse=True):
if lemma.startswith(p) and len(lemma) > len(p) + 2 \
and not lemma.startswith(_INSEP):
return p
return None
def finite(lemma, tense, person, number):
"""Present/past finite. For separable verbs the returned string is the
UniMorph SEPARATED form 'stem prefix' (realizer places prefix per V2)."""
slot = "prs" if tense == "present" else "pst"
if lemma in _CORE and _CORE[lemma].get(slot, {}).get((person, number)):
return _CORE[lemma][slot][(person, number)], "lexicon"
rec = _VERBS.get(lemma)
if rec and rec.get(slot, {}).get((person, number)):
return rec[slot][(person, number)], "lexicon"
# rule fallback (present only reliable; past weak -te)
stem = _stem(lemma)
pref = separable_prefix(lemma)
if pref:
stem = _stem(lemma[len(pref):])
if tense == "present":
end = _PRS_ENDINGS[(person, number)]
if stem.endswith(("t", "d", "chn", "ffn", "gn")) and end in ("st", "t"):
end = "e" + end
form = stem + end
else:
form = stem + ("ete" if stem.endswith(("t", "d")) else "te")
if (person, number) == ("second", "singular"):
form += "st"
elif number == "plural" and person != "second":
form += "n"
elif (person, number) == ("second", "plural"):
form += "t"
if pref:
return f"{form} {pref}", "rule"
return form, "rule"
def _weak_t(stem):
return stem + ("et" if stem.endswith(("t", "d", "chn", "ffn", "gn")) else "t")
def past_participle(lemma):
if lemma in _CORE:
return _CORE[lemma]["ppart"], "lexicon"
rec = _VERBS.get(lemma)
if rec and rec.get("ppart"):
return rec["ppart"], "lexicon"
stem = _stem(lemma)
pref = separable_prefix(lemma)
_INSEP = ("be", "ge", "er", "ver", "zer", "ent", "emp", "miss")
if pref:
inner = _stem(lemma[len(pref):])
return pref + "ge" + _weak_t(inner), "rule"
if lemma.startswith(_INSEP):
return _weak_t(stem), "rule"
return "ge" + _weak_t(stem), "rule"
def nonfinite(lemma, req):
if req == "ppart":
return past_participle(lemma)
return lemma, "lexicon" if lemma in _VERBS else "rule" # infinitive
def perfect_aux(lemma):
return "sein" if lemma in _SEIN else "haben"
# ── adjectives ────────────────────────────────────────────────────────────────---
_ADJ_IRREG_SPRL = {"gut": "best", "groß": "größt", "hoch": "höchst",
"nah": "nächst", "viel": "meist", "gern": "liebst"}
def comparative(lemma):
rec = _ADJS.get(lemma)
if rec and rec.get("cmpr"):
return rec["cmpr"], "lexicon"
return lemma + "er", "rule"
def superlative(lemma):
"""Return the bare superlative STEM (realizer adds 'am ...en' or '-e' ending)."""
if lemma in _ADJ_IRREG_SPRL:
return _ADJ_IRREG_SPRL[lemma], "lexicon"
# derive from the comparative so umlaut is carried (alt->älter->ältest)
cmpr, cconf = comparative(lemma)
base = cmpr[:-2] if cmpr.endswith("er") else lemma
end = "est" if base.endswith(("t", "d", "s", "ß", "z", "sch")) else "st"
return base + end, cconf
def lexicon_stats():
return {
"source": "UniMorph deu (primary) + kaikki.org German (gap-fill gender/plural)",
"license": "CC-BY-SA 3.0 (UniMorph); CC-BY-SA/GFDL (Wiktionary)",
"noun_lemmas": len(_NOUNS),
"nouns_with_gender": sum(1 for v in _NOUNS.values() if v.get("g")),
"nouns_with_plural": sum(1 for v in _NOUNS.values() if v.get("pl")),
"verb_lemmas": len(_VERBS),
"verbs_with_ppart": sum(1 for v in _VERBS.values() if v.get("ppart")),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
for w in ("Hund", "Frau", "Kind", "Mann", "Buch", "Blume"):
print(f" {w}: gender={noun_gender(w)} pl={pluralize(w)} "
f"gen.sg={decline_noun(w, 'GEN', 'singular')} "
f"dat.pl={decline_noun(w, 'DAT', 'plural')}")
for v in ("machen", "gehen", "aufstehen", "sein", "haben", "arbeiten"):
print(f" {v}: 3sg.prs={finite(v, 'present', 'third', 'singular')} "
f"3sg.pst={finite(v, 'past', 'third', 'singular')} "
f"ppart={past_participle(v)} aux={perfect_aux(v)} sep={separable_prefix(v)}")
for a in ("schnell", "gut", "groß", "alt"):
print(f" {a}: cmpr={comparative(a)} sprl={superlative(a)}")
-562
View File
@@ -1,562 +0,0 @@
"""morphology_es_full.py — production-grade Spanish morphological generator.
NOT a toy. Backed by a real, broad, licensed lexicon:
UniMorph Spanish (github.com/unimorph/spa, CC-BY-SA 3.0, Wiktionary-derived)
1,196,245 inflected forms:
6,695 verb lemmas full paradigms: indicative (present/preterite/
imperfect/future), conditional, present & imperfect
subjunctive, affirmative imperative, formal/informal
48,353 noun lemmas WITH inherent gender (N;FEM/MASC;SG/PL)
16,984 adj lemmas gender + number paradigms
Fallbacks (so we degrade, never crash, on out-of-vocabulary input):
- verbs : mlconjug3 (ML paradigm model, conjugates ANY Spanish verb) then a
hand-rolled regular-ending generator
- nouns : gender heuristic (endings) + regular pluralization
- adjs : -o/-a gender rule + regular pluralization
Every generated form carries a CONFIDENCE flag:
"lexicon" form came straight from UniMorph (trust: high)
"model" form came from mlconjug3 (trust: high)
"rule" form came from a deterministic rule (trust: medium)
"fallback" we could not inflect; returned lemma as-is (trust: low FLAG)
Public API (used by realizer_es.py):
conjugate(lemma, mood, tense, person, number, formality="informal") -> (form, conf)
participle(lemma) -> (form, conf) # past participle (compound tenses)
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"
inflect_noun(lemma, number) -> (form, conf)
inflect_adj(lemma, gender, number) -> (form, conf)
attach_enclitics(verb_form, clitics) -> str # accent-correct enclisis
lexicon_stats() -> dict
"""
import os
import pickle
import unicodedata
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "spa.unimorph")
_CACHE = os.path.join(_HERE, "data", "es_morph_cache.pkl")
# ── canonical feature keys the realizer speaks, mapped to UniMorph tags ─────────
# mood/tense pair -> the UniMorph feature substring that identifies it
_VERB_KEYMAP = {
("ind", "present"): ("IND", "PRS", None),
("ind", "preterite"): ("IND", "PST", "PFV"),
("ind", "imperfect"): ("IND", "PST", "IPFV"),
("ind", "future"): ("IND", "FUT", None),
("ind", "conditional"):("COND", None, None),
("sbjv", "present"): ("SBJV", "PRS", None),
("sbjv", "imperfect"): ("SBJV", "PST", "LGSPEC1"), # -ra form
("imp", "present"): ("POS", "IMP", None),
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
# ── build / load the compact lexicon ───────────────────────────────────────────
def _feat_set(tag):
return set(tag.split(";"))
def _build_cache():
verbs = {} # (lemma, canonkey) -> form canonkey e.g. "ind|present|1|SG|infm"
nouns = {} # lemma -> {"g": "m"/"f", "SG": form, "PL": form}
adjs = {} # lemma -> {("m","SG"): form, ...}
part = {} # lemma -> masc-sg participle
ger = {} # lemma -> gerund
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V":
# skip clitic-bearing rows (we generate clitics ourselves)
if "PRO" in f:
continue
if "V.PTCP" in f and "PST" in f and "MASC" in f and "SG" in f:
part.setdefault(lemma, form)
continue
if "V.CVB" in f or "NFIN" in f or "V.PTCP" in f:
if "V.CVB" in f:
ger.setdefault(lemma, form)
continue
# identify mood/tense
mt = None
for (mood, tense), (a, b, c) in _VERB_KEYMAP.items():
if a not in f:
continue
if b is not None and b not in f:
continue
if c is not None and c not in f:
continue
# disambiguate IND;PST needing PFV vs IPFV
if a == "IND" and b == "PST" and c not in f:
continue
mt = (mood, tense)
break
if mt is None:
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
formal = "form" if "FORM" in f else ("infm" if "INFM" in f else "any")
key = f"{mt[0]}|{mt[1]}|{person}|{number}|{formal}"
verbs.setdefault((lemma, key), form)
elif head == "N":
# substring test handles epicene "MASC+FEM" (-> masc citation)
g = "m" if "MASC" in tag else ("f" if "FEM" in tag else None)
num = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if num is None:
continue
# store forms keyed by (gender,number); animate nouns list BOTH
# genders under one lemma (niño -> niño/niña). Resolve citation
# gender in a post-pass (gender of the row whose form == lemma).
d = nouns.setdefault(lemma, {})
d.setdefault("_rows", []).append((g, num, form))
elif head == "ADJ":
g = "m" if "MASC" in tag else ("f" if "FEM" in tag else "m")
num = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if num is None:
continue
adjs.setdefault(lemma, {})[(g, num)] = form
# post-pass: resolve noun citation gender + default SG/PL forms
for lemma, d in nouns.items():
rows = d.pop("_rows", [])
# citation gender = gender of the row whose form == lemma; else first MASC;
# else first seen gender.
cite_g = None
for g, num, form in rows:
if form == lemma and g:
cite_g = g
break
if cite_g is None:
for g, num, form in rows:
if g == "m":
cite_g = "m"
break
if cite_g is None:
cite_g = next((g for g, _, _ in rows if g), "m")
d["g"] = cite_g
for g, num, form in rows:
d[(g, num)] = form
d["SG"] = d.get((cite_g, "SG")) or next((f for g, n, f in rows if n == "SG"), lemma)
d["PL"] = d.get((cite_g, "PL")) or next((f for g, n, f in rows if n == "PL"), None)
# post-pass: UniMorph omits the identity inflection (masc-sg == lemma) for
# adjectives, so fill it in; without this a fem-sg row wrongly satisfies a
# masc-sg request (alto -> alta bug).
for lemma, d in adjs.items():
d.setdefault(("m", "SG"), lemma)
data = {"verbs": verbs, "nouns": nouns, "adjs": adjs, "part": part, "ger": ger}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE) and os.path.getmtime(_CACHE) >= os.path.getmtime(_UNIMORPH):
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _NOUNS, _ADJS, _PART, _GER = (
_LEX["verbs"], _LEX["nouns"], _LEX["adjs"], _LEX["part"], _LEX["ger"])
# ── mlconjug3 fallback (lazy) ───────────────────────────────────────────────────
_MLC = None
_MLC_TENSE = { # (mood,tense) -> (mlconjug mood label, tense label)
("ind", "present"): ("Indicativo", "Indicativo presente"),
("ind", "preterite"): ("Indicativo", "Indicativo pretérito perfecto simple"),
("ind", "imperfect"): ("Indicativo", "Indicativo pretérito imperfecto"),
("ind", "future"): ("Indicativo", "Indicativo futuro"),
("ind", "conditional"): ("Condicional", "Condicional Condicional"),
("sbjv", "present"): ("Subjuntivo", "Subjuntivo presente"),
("sbjv", "imperfect"): ("Subjuntivo", "Subjuntivo pretérito imperfecto 1"),
("imp", "present"): ("Imperativo", "Imperativo Afirmativo"),
}
_MLC_SLOT = { # (person,number) -> mlconjug slot key
("first", "singular"): "1s", ("second", "singular"): "2s",
("third", "singular"): "3s", ("first", "plural"): "1p",
("second", "plural"): "2p", ("third", "plural"): "3p",
}
def _mlc_conjugate(lemma, mood, tense, person, number):
global _MLC
try:
if _MLC is None:
from mlconjug3 import Conjugator
_MLC = Conjugator(language="es")
v = _MLC.conjugate(lemma)
if v is None:
return None
info = v.conjug_info
m, t = _MLC_TENSE.get((mood, tense), (None, None))
if m is None or m not in info or t not in info[m]:
return None
block = info[m][t]
slot = _MLC_SLOT.get((person, number))
if isinstance(block, dict) and slot in block and block[slot]:
return block[slot]
return None
except Exception:
return None
# ── regular-ending rule fallback (last resort, deterministic) ───────────────────
def _vclass(lemma):
return lemma[-2:] if lemma[-2:] in ("ar", "er", "ir") else "ar"
def _stem(lemma):
return lemma[:-2]
_REG = {
("ind", "present", "ar"): ["o", "as", "a", "amos", "áis", "an"],
("ind", "present", "er"): ["o", "es", "e", "emos", "éis", "en"],
("ind", "present", "ir"): ["o", "es", "e", "imos", "ís", "en"],
("ind", "preterite", "ar"): ["é", "aste", "ó", "amos", "asteis", "aron"],
("ind", "preterite", "er"): ["í", "iste", "", "imos", "isteis", "ieron"],
("ind", "preterite", "ir"): ["í", "iste", "", "imos", "isteis", "ieron"],
("ind", "imperfect", "ar"): ["aba", "abas", "aba", "ábamos", "abais", "aban"],
("ind", "imperfect", "er"): ["ía", "ías", "ía", "íamos", "íais", "ían"],
("ind", "imperfect", "ir"): ["ía", "ías", "ía", "íamos", "íais", "ían"],
("sbjv", "present", "ar"): ["e", "es", "e", "emos", "éis", "en"],
("sbjv", "present", "er"): ["a", "as", "a", "amos", "áis", "an"],
("sbjv", "present", "ir"): ["a", "as", "a", "amos", "áis", "an"],
("sbjv", "imperfect", "ar"): ["ara", "aras", "ara", "áramos", "arais", "aran"],
("sbjv", "imperfect", "er"): ["iera", "ieras", "iera", "iéramos", "ierais", "ieran"],
("sbjv", "imperfect", "ir"): ["iera", "ieras", "iera", "iéramos", "ierais", "ieran"],
}
_FUT = ["é", "ás", "á", "emos", "éis", "án"]
_COND = ["ía", "ías", "ía", "íamos", "íais", "ían"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _rule_conjugate(lemma, mood, tense, person, number):
if len(lemma) < 3 or lemma[-2:] not in ("ar", "er", "ir"):
return None
vc, st, i = _vclass(lemma), _stem(lemma), _slot_idx(person, number)
if tense == "future":
return lemma + _FUT[i]
if tense == "conditional":
return lemma + _COND[i]
table = _REG.get((mood, tense, vc))
if table:
return st + table[i]
if mood == "imp" and tense == "present":
# affirmative tú imperative = 3sg present indicative
pres = _REG.get(("ind", "present", vc))
return st + pres[2] if number == "singular" else st + pres[5]
return None
# ── PUBLIC: verb conjugation ────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number, formality="informal"):
"""Return (surface, confidence). mood in ind|sbjv|imp; tense per _VERB_KEYMAP."""
lemma = lemma.strip().lower()
p, n = _PERSON.get(person), _NUMBER.get(number)
formal = "form" if formality == "formal" else "infm"
if p and n:
for fkey in (formal, "any", "infm" if formal == "form" else "form"):
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}|{fkey}"))
if form:
return form, "lexicon"
m = _mlc_conjugate(lemma, mood, tense, person, number)
if m:
return m, "model"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r:
return r, "rule"
return lemma, "fallback"
_IRREG_PART = { # guarantee the common irregular participles
"escribir": "escrito", "describir": "descrito", "abrir": "abierto",
"cubrir": "cubierto", "descubrir": "descubierto", "morir": "muerto",
"poner": "puesto", "ver": "visto", "volver": "vuelto", "devolver": "devuelto",
"hacer": "hecho", "deshacer": "deshecho", "decir": "dicho", "romper": "roto",
"resolver": "resuelto", "freír": "frito", "imprimir": "impreso",
"satisfacer": "satisfecho", "prever": "previsto", "revolver": "revuelto",
}
def participle(lemma):
lemma = lemma.strip().lower()
if lemma in _IRREG_PART:
return _IRREG_PART[lemma], "lexicon"
if lemma in _PART:
return _PART[lemma], "lexicon"
if lemma.endswith("ar"):
return lemma[:-2] + "ado", "rule"
if lemma[-2:] in ("er", "ir"):
return lemma[:-2] + "ido", "rule"
return lemma, "fallback"
_IRREG_GER = {"dormir": "durmiendo", "morir": "muriendo", "pedir": "pidiendo",
"sentir": "sintiendo", "mentir": "mintiendo", "servir": "sirviendo",
"venir": "viniendo", "decir": "diciendo", "poder": "pudiendo",
"ir": "yendo", "leer": "leyendo", "creer": "creyendo",
"oír": "oyendo", "traer": "trayendo", "caer": "cayendo",
"construir": "construyendo", "huir": "huyendo", "reír": "riendo"}
def gerund(lemma):
lemma = lemma.strip().lower()
if lemma in _IRREG_GER:
return _IRREG_GER[lemma], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
if lemma.endswith("ar"):
return lemma[:-2] + "ando", "rule"
if lemma[-2:] in ("er", "ir"):
return lemma[:-2] + "iendo", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ────────────────────────────────────────────────
_INVARIANT_PL = {"lunes", "martes", "miércoles", "jueves", "viernes",
"crisis", "tesis", "análisis", "dosis", "virus", "paraguas"}
def _gender_heuristic(noun):
for suf, g in (("ión", "f"), ("dad", "f"), ("tad", "f"), ("umbre", "f"),
("sis", "f"), ("ez", "f"), ("triz", "f"),
("ema", "m"), ("ama", "m"), ("oma", "m"), ("aje", "m"),
("or", "m"), ("án", "m"), ("ín", "m")):
if noun.endswith(suf):
return g
if noun.endswith("o"):
return "m"
if noun.endswith("a"):
return "f"
return "m"
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g"):
return d["g"]
return _gender_heuristic(lemma)
def _regular_plural(noun):
if noun in _INVARIANT_PL:
return noun
if not noun:
return noun
last = noun[-1]
if last == "z":
return noun[:-1] + "ces"
if last in "aeiouáéíóú":
# stressed final vowel í/ú -> +es (rubí->rubíes), else +s
if last in "íú":
return noun + "es"
return noun + "s"
if last == "s":
# esdrújula / stress-final handled crudely; most polysyllables invariant
return noun
return noun + "es"
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
num = "SG" if number == "singular" else "PL"
if d:
# honor a requested gender for animate nouns (gato -> gata)
if gender and (gender, num) in d:
return d[(gender, num)], "lexicon"
if d.get(num):
return d[num], "lexicon"
if number == "singular":
return lemma, "rule" if not d else "lexicon"
return _regular_plural(lemma), "rule"
# ── PUBLIC: adjective agreement ─────────────────────────────────────────────────
_INV_GENDER_ADJ = {"español": "española", "trabajador": "trabajadora",
"hablador": "habladora", "encantador": "encantadora",
"alemán": "alemana", "francés": "francesa", "inglés": "inglesa"}
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
d = _ADJS.get(lemma)
num = "SG" if number == "singular" else "PL"
if d:
form = d.get((gender, num))
if form:
return form, "lexicon"
# gender-invariant adjective (grande, feliz, azul): fem == masc.
# For a missing plural, pluralize this gender's singular form.
sg = d.get((gender, "SG")) or d.get(("m", "SG")) or lemma
if number == "plural":
return _regular_plural(sg), "rule"
return sg, "lexicon"
# rule fallback
a = lemma
if gender == "f":
if a in _INV_GENDER_ADJ:
a = _INV_GENDER_ADJ[a]
elif a.endswith("o"):
a = a[:-1] + "a"
if number == "plural":
a = _regular_plural(a)
return a, ("rule" if (a != lemma or gender == "m") else "rule")
# ── PUBLIC: clitic enclisis (dá + me + lo -> dámelo) ────────────────────────────
def _strip_accents(s):
return "".join(c for c in unicodedata.normalize("NFD", s)
if unicodedata.category(c) != "Mn")
def _count_syllables_vowelgroups(word):
# crude: count vowel groups
w = _strip_accents(word).lower()
groups, prev = 0, False
for ch in w:
isv = ch in "aeiou"
if isv and not prev:
groups += 1
prev = isv
return groups
def _host_stress_from_end(word):
"""Stressed-syllable index counted from the end (1=last) of a verb host."""
syls = _count_syllables_vowelgroups(word)
if any(c in "áéíóú" for c in word):
return None # already carries its own accent
if word[-2:] in ("ar", "er", "ir"): # infinitive: oxytone
return 1
if word.endswith("ndo"): # gerund: paroxytone
return 2
if word[-1:] in "aeiouns" and syls >= 2: # default paroxytone
return 2
return 1 # monosyllable / consonant-final oxytone
def attach_enclitics(verb_form, clitics):
"""Append clitic pronouns to a verb (imperative/infinitive/gerund enclisis)
and add a written accent when the resulting word becomes esdrújula/
sobreesdrújula (stress >= 3 syllables from the end): +me+lo -> dámelo,
lleva+me -> llévame, but dar+te -> darte and da+me -> dame (no accent)."""
if not clitics:
return verb_form
tail = "".join(clitics)
if any(c in "áéíóú" for c in verb_form): # host already accented
return verb_form + tail
sfe = _host_stress_from_end(verb_form)
total_sfe = sfe + len(clitics) # each clitic = 1 syllable
if total_sfe >= 3:
return _accentuate_nucleus(verb_form, sfe) + tail
return verb_form + tail
def _accentuate_nucleus(word, sfe):
"""Put a written accent on the syllable `sfe` positions from the word's end."""
vowels = "aeiou"
nuclei = [i for i, ch in enumerate(word) if ch in vowels]
if not nuclei or sfe > len(nuclei):
return word
i = nuclei[-sfe]
acc = {"a": "á", "e": "é", "i": "í", "o": "ó", "u": "ú"}
return word[:i] + acc[word[i]] + word[i + 1:]
def _accentuate_last_stressed(word):
# Restore the host's ORIGINAL lexical stress with a written accent.
# Default Spanish stress: word ending in vowel/n/s -> penultimate syllable;
# otherwise (e.g. infinitives in -r) -> last syllable.
vowels = "aeiou"
nuclei = [i for i, ch in enumerate(word) if ch in vowels]
if not nuclei:
return word
if word[-1] in "aeiouns" and len(nuclei) >= 2:
i = nuclei[-2] # paroxytone: penult nucleus
else:
i = nuclei[-1] # oxytone / monosyllable: last nucleus
acc = {"a": "á", "e": "é", "i": "í", "o": "ó", "u": "ú"}
return word[:i] + acc[word[i]] + word[i + 1:]
def lexicon_stats():
return {
"source": "UniMorph Spanish (github.com/unimorph/spa)",
"license": "CC-BY-SA 3.0 (Wiktionary-derived)",
"total_forms": sum(len(v) for v in (_VERBS, _NOUNS, _ADJS)) if False else None,
"verb_forms": len(_VERBS),
"verb_lemmas": len({k[0] for k in _VERBS}),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
"participles": len(_PART),
"gerunds": len(_GER),
}
if __name__ == "__main__":
import json
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("hablar", "ind", "present", "first", "singular", "hablo"),
("comer", "ind", "present", "third", "plural", "comen"),
("vivir", "ind", "present", "first", "plural", "vivimos"),
("ser", "ind", "present", "third", "singular", "es"),
("ir", "ind", "preterite", "first", "singular", "fui"),
("tener", "ind", "future", "first", "singular", "tendré"),
("hacer", "sbjv", "present", "first", "singular", "haga"),
("dormir", "ind", "present", "first", "singular", "duermo"),
("pensar", "sbjv", "present", "third", "singular", "piense"),
("dar", "ind", "preterite", "third", "singular", "dio"),
("poner", "ind", "conditional", "first", "singular", "pondría"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
if got == exp:
ok += 1
print(f" {flag}{lemma:8} {mood}/{tense} {per[:3]}.{num[:2]:3} -> {got:14} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" gender casa:", noun_gender("casa"), "| problema:", noun_gender("problema"),
"| agua:", noun_gender("agua"), "| mano:", noun_gender("mano"))
print(" plural: luz->", inflect_noun("luz", "plural"), "| rey->", inflect_noun("rey", "plural"))
print(" adj: rojo/f/pl->", inflect_adj("rojo", "f", "plural"),
"| feliz/m/pl->", inflect_adj("feliz", "m", "plural"),
"| grande/f/pl->", inflect_adj("grande", "f", "plural"))
print(" enclisis: da+[me,lo]->", attach_enclitics("da", ["me", "lo"]),
"| di+[me]->", attach_enclitics("di", ["me"]),
"| dar+[se,lo]->", attach_enclitics("dar", ["se", "lo"]))
-629
View File
@@ -1,629 +0,0 @@
"""morphology_fr_full.py — production-grade French morphological generator.
Same architecture as morphology_it_full.py (shared Romance engine); French-specific
data and rules swapped in. Backed by three real, Wiktionary-lineage sources:
VERBS
UniMorph French (github.com/unimorph/fra, CC-BY-SA 3.0)
7,535 verb lemmas × full paradigm, CLEAN orthography:
indicatif présent / imparfait (PST;IPFV) / passé simple (PST;PFV) /
futur, conditionnel (COND), subjonctif présent (SBJV;PRS) /
subjonctif imparfait (SBJV;PST), impératif (POS;IMP), infinitif (NFIN),
participe présent (V.CVB/V.PTCP;PRS), participe passé (V.PTCP;PST, m.sg).
fr_irreg_verbs.json high-frequency verbs UniMorph MISSES or mis-slots,
above all ÊTRE (absent from UniMorph fra), plus avoir/aller/faire/ the
auxiliaries the passé-composé + être-agreement system depends on. Extracted
from kaikki.org French (build_fr_irreg.py), reflexive/multiword forms
dropped. This layer takes PRIORITY.
NOUNS + ADJECTIVES kaikki.org French (Wiktionary extract, CC-BY-SA 3.0)
noun lemmas WITH inherent gender (head-template arg) + real plural
(cheval->chevaux, œil->yeux, invariable -s/-x/-z), resolved PER LEMMA.
adjective lemmas with real feminine + plural (petit->petite/petits/petites,
beau->belle/beaux/belles, heureux->heureuse, rouge invariant-gender).
Fallbacks (degrade, never crash, on OOV input):
verbs : rule generator for -er / -ir(-iss-) / -re (with -cer/-ger spelling,
future/conditional stems, imparfait/subjonctif endings)
nouns : gender heuristic (endings) + rule pluralization (-al->-aux, -eau->-eaux)
adjs : fem/plural agreement rules (-er->-ère, -eux->-euse, -f->-ve, +e default)
Confidence flag on every form: "lexicon" | "rule" | "fallback".
Public API (used by realizer_fr.py): identical signature to morphology_it_full.
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "fra.unimorph")
_IRREG = os.path.join(_HERE, "data", "fr_irreg_verbs.json")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_fr.jsonl")
_CACHE = os.path.join(_HERE, "data", "fr_morph_cache.pkl")
# ── (mood, tense) -> UniMorph feature set that must ALL be present ────────────────
_VERB_KEYMAP = {
("ind", "present"): {"IND", "PRS"},
("ind", "imperfect"): {"IND", "PST", "IPFV"}, # imparfait
("ind", "passe_simple"): {"IND", "PST", "PFV"}, # passé simple
("ind", "future"): {"IND", "FUT"},
("ind", "conditional"): {"COND"}, # French: V;COND;1;SG
("sbjv", "present"): {"SBJV", "PRS"},
("sbjv", "imperfect"): {"SBJV", "PST"},
("imp", "affirmative"): {"POS", "IMP"},
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── build verb lexicon from UniMorph ─────────────────────────────────────────────
def _build_verbs():
verbs = {}
part = {}
ger = {}
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f:
part.setdefault(lemma, form)
elif "PRS" in f:
ger.setdefault(lemma, form)
continue
if head == "V.CVB":
if "PRS" in f:
ger.setdefault(lemma, form)
continue
if head != "V":
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
for (mood, tense), req in _VERB_KEYMAP.items():
if not req <= f:
continue
if tense == "imperfect" and "PFV" in f:
continue
if tense == "passe_simple" and "IPFV" in f:
continue
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), form)
break
return verbs, part, ger
# ── kaikki nouns + adjectives ────────────────────────────────────────────────────
_EXCL_FORM_TAGS = {"alternative", "archaic", "obsolete", "dialectal", "regional",
"diminutive", "augmentative", "pejorative", "comparative",
"superlative", "misspelling", "rare", "informal", "literary",
"poetic", "error-unrecognized-form", "construed", "collective",
"nonstandard", "dated", "Louisiana", "Switzerland", "Belgium"}
def _kaikki_gender(arg):
if not arg:
return None
a = str(arg).lower()
if a.startswith("f"):
return "f"
if a.startswith("m"):
return "m"
return None
def _build_nouns_adjs():
nouns = {}
adjs = {}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
pos = d.get("pos")
word = d.get("word", "")
if not word or " " in word:
continue
forms = d.get("forms", []) or []
if pos == "noun":
ht = d.get("head_templates") or []
g = None
if ht:
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
if g is None:
tags = d.get("tags") or []
if "feminine" in tags:
g = "f"
elif "masculine" in tags:
g = "m"
pl = None
for x in forms:
t = set(x.get("tags") or [])
if "plural" in t and not (t & _EXCL_FORM_TAGS):
fm = x.get("form")
if fm and " " not in fm and fm not in ("#", "-", ""):
pl = fm
break
if word not in nouns:
nouns[word] = {"g": g, "SG": word, "PL": pl}
else:
cur = nouns[word]
if cur.get("g") is None and g:
cur["g"] = g
if not cur.get("PL") and pl:
cur["PL"] = pl
elif pos == "adj":
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word)
for x in forms:
t = set(x.get("tags") or [])
fm = x.get("form")
if not fm or " " in fm or (t & _EXCL_FORM_TAGS):
continue
if "feminine" in t and "plural" in t:
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
elif "masculine" in t and "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
elif "feminine" in t:
d0[("f", "SG")] = d0.get(("f", "SG")) or fm
elif "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
return nouns, adjs
def _build_cache():
verbs, part, ger = _build_verbs()
nouns, adjs = _build_nouns_adjs()
with open(_IRREG, encoding="utf-8") as fh:
irreg = json.load(fh)
data = {"verbs": verbs, "part": part, "ger": ger,
"nouns": nouns, "adjs": adjs, "irreg": irreg}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [_UNIMORPH, _KAIKKI, _IRREG]
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PART, _GER, _NOUNS, _ADJS, _IRREGV = (
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"],
_LEX["irreg"])
# ── regular-ending rule fallback ─────────────────────────────────────────────────
def _vclass(lemma):
if lemma.endswith("er"):
return "er"
if lemma.endswith("ir"):
return "ir"
if lemma.endswith("re"):
return "re"
if lemma.endswith("oir"):
return "oir"
return None
# present-tense endings [1sg,2sg,3sg,1pl,2pl,3pl]
_REG_PRES = {
"er": ["e", "es", "e", "ons", "ez", "ent"],
"ir": ["is", "is", "it", "issons", "issez", "issent"], # -iss- class (finir)
"re": ["s", "s", "", "ons", "ez", "ent"], # vendre: vends/vend
}
_REG_IMPF = ["ais", "ais", "ait", "ions", "iez", "aient"] # attaches to pres-1pl stem
_REG_SUBJ = ["e", "es", "e", "ions", "iez", "ent"] # attaches to 3pl stem
_REG_PS = { # passé simple
"er": ["ai", "as", "a", "âmes", "âtes", "èrent"],
"ir": ["is", "is", "it", "îmes", "îtes", "irent"],
"re": ["is", "is", "it", "îmes", "îtes", "irent"],
}
_FUT = ["ai", "as", "a", "ons", "ez", "ont"]
_COND = ["ais", "ais", "ait", "ions", "iez", "aient"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _fut_stem(lemma, vc):
"""Future/conditional stem = infinitive (drop final -e of -re)."""
if vc == "re":
return lemma[:-1] # vendre -> vendr-
return lemma # parler-, finir-
def _pres_1pl_stem(lemma, vc):
"""Imparfait stem = present 1pl minus -ons (parlons->parl-, finissons->finiss-)."""
if vc == "er":
stem = lemma[:-2]
if stem.endswith("g"):
return stem + "e" # mangeons -> mange- (imparfait mangeais)
if stem.endswith("c"):
return stem[:-1] + "ç" # commençons -> commenç-
return stem
if vc == "ir":
return lemma[:-1] + "iss" # finir -> finiss-
if vc == "re":
return lemma[:-2] # vendre -> vend-
return lemma[:-2]
def _apply_er_spelling(stem, ending):
"""-cer/-ger softening before a/o (commençons, mangeons)."""
if ending and ending[0] in ("a", "o"):
if stem.endswith("c"):
return stem[:-1] + "ç" + ending
if stem.endswith("g"):
return stem + "e" + ending
return stem + ending
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
i = _slot_idx(person, number)
if mood == "ind" and tense in ("future", "conditional"):
stem = _fut_stem(lemma, vc)
end = (_FUT if tense == "future" else _COND)[i]
return stem + end
if mood == "ind" and tense == "present":
table = _REG_PRES.get("ir" if vc == "ir" else vc)
if not table:
return None
body = lemma[:-2] if vc in ("er", "re") else lemma[:-1] if vc == "ir" else lemma[:-2]
if vc == "ir":
body = lemma[:-2] # fin- ; endings carry -iss-
end = table[i]
return body + end
end = table[i]
if vc == "er":
return _apply_er_spelling(body, end)
return body + end
if mood == "ind" and tense == "imperfect":
stem = _pres_1pl_stem(lemma, vc)
return stem + _REG_IMPF[i]
if mood == "ind" and tense == "passe_simple":
table = _REG_PS.get("ir" if vc == "ir" else vc)
if not table:
return None
body = lemma[:-2] if vc in ("er", "re") else lemma[:-2]
end = table[i]
if vc == "er":
return _apply_er_spelling(body, end)
return body + end
if mood == "sbjv" and tense == "present":
# subjonctif: present-3pl stem + e/es/e/ions/iez/ent
stem3 = _pres_1pl_stem(lemma, vc) if vc == "ir" else (
lemma[:-2] if vc in ("er", "re") else lemma[:-2])
if vc == "ir":
stem3 = lemma[:-2] + "iss"
end = _REG_SUBJ[i]
if vc == "er":
return _apply_er_spelling(stem3, end)
return stem3 + end
if mood == "imp" and tense == "affirmative":
# impératif ~ present indicative (tu drops -s for -er verbs)
pres = _rule_conjugate(lemma, "ind", "present", person, number)
if pres and vc == "er" and person == "second" and number == "singular":
return pres[:-1] if pres.endswith("es") else pres
return pres
return None
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
"""Return (surface, confidence)."""
lemma = lemma.strip().lower()
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{number}"
ir = _IRREGV.get(lemma)
if ir and key in ir:
return ir[key], "lexicon"
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
if form:
return form, "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r:
return r, "rule"
return lemma, "fallback"
# ── PUBLIC: participle + gerund/participe présent ────────────────────────────────
def _participle_msg(lemma):
ir = _IRREGV.get(lemma)
if ir and "part" in ir:
return ir["part"], "lexicon"
if lemma in _PART:
return _PART[lemma], "lexicon"
return None, None
# irregular participle fem/plural quirks (drop circonflexe: dû->due, dus)
_PART_FIX = {"": {"f|SG": "due", "m|PL": "dus", "f|PL": "dues"}}
def participle(lemma, gender="m", number="singular"):
"""Past participle with French gender/number agreement.
m.sg = base; f.sg = base+e; m.pl = base+s (invariable if base ends s/x);
f.pl = f.sg+s."""
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
msg, src = _participle_msg(lemma)
conf = "lexicon"
if msg is None:
vc = _vclass(lemma)
if vc == "er":
msg = lemma[:-2] + "é"
elif vc == "ir":
msg = lemma[:-1] # finir -> fini, partir -> parti
elif vc == "re":
msg = lemma[:-2] + "u" # vendre -> vendu
elif vc == "oir":
msg = lemma[:-3] + "u" # (rough) recevoir handled by irreg
else:
return lemma, "fallback"
conf = "rule"
fix = _PART_FIX.get(msg)
if fix and f"{g}|{num}" in fix:
return fix[f"{g}|{num}"], conf
if g == "m" and num == "SG":
return msg, conf
fem = msg + "e" if not msg.endswith("e") else msg
if g == "f" and num == "SG":
return fem, conf
if g == "m" and num == "PL":
return msg if msg.endswith(("s", "x")) else msg + "s", conf
# f|PL
return fem + "s", conf
def gerund(lemma):
"""Participe présent (base for gérondif 'en -ant')."""
lemma = lemma.strip().lower()
ir = _IRREGV.get(lemma)
if ir and "ger" in ir:
return ir["ger"], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
vc = _vclass(lemma)
if vc == "er":
stem = lemma[:-2]
if stem.endswith("g"):
return stem + "eant", "rule"
if stem.endswith("c"):
return stem[:-1] + "çant", "rule"
return stem + "ant", "rule"
if vc == "ir":
return lemma[:-2] + "issant", "rule"
if vc == "re":
return lemma[:-2] + "ant", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
_FEM_SUF = ("tion", "sion", "aison", "ance", "ence", "ette", "elle", "esse",
"ude", "ade", "ée", "", "tié", "ie", "ise", "ure", "eur")
_MASC_SUF = ("ment", "age", "eau", "isme", "oir", "ier", "eur", "in", "on")
def _gender_heuristic(noun):
for suf in _FEM_SUF:
if noun.endswith(suf):
return "f"
for suf in _MASC_SUF:
if noun.endswith(suf):
return "m"
if noun.endswith("e"):
return "f"
return "m"
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g") in ("m", "f"):
return d["g"]
return _gender_heuristic(lemma)
# closed sets for French plural irregularities
_OU_X = {"bijou", "caillou", "chou", "genou", "hibou", "joujou", "pou"}
_AIL_AUX = {"travail", "vitrail", "corail", "émail", "bail", "soupirail", "vantail"}
_AL_S = {"bal", "carnaval", "festival", "récital", "chacal", "régal", "cal", "aval"}
def _rule_plural(noun, gender):
"""Deterministic French pluralization. (form, ok); ok=False FLAGS ambiguity."""
if not noun:
return noun, True
if noun[-1:] in ("s", "x", "z"):
return noun, True # invariable
if noun in _OU_X:
return noun + "x", True
if noun.endswith(("eau", "au", "eu")):
if noun in ("pneu", "bleu", "landau", "sarrau"):
return noun + "s", True
return noun + "x", True # bateau->bateaux, jeu->jeux
if noun.endswith("al"):
if noun in _AL_S:
return noun + "s", True
return noun[:-2] + "aux", True # cheval->chevaux
if noun.endswith("ail"):
if noun in _AIL_AUX:
return noun[:-3] + "aux", True # travail->travaux
return noun + "s", True
return noun + "s", True # default
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if number == "singular":
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
if d and d.get("PL"):
return d["PL"], "lexicon"
g = gender or noun_gender(lemma)
form, ok = _rule_plural(lemma, g)
return form, ("rule" if ok else "fallback")
# adjectives whose kaikki entries are unreliable: audited forms
_ADJ_FIX = {
"beau": {("m", "SG"): "beau", ("f", "SG"): "belle",
("m", "PL"): "beaux", ("f", "PL"): "belles"},
"nouveau": {("m", "SG"): "nouveau", ("f", "SG"): "nouvelle",
("m", "PL"): "nouveaux", ("f", "PL"): "nouvelles"},
"vieux": {("m", "SG"): "vieux", ("f", "SG"): "vieille",
("m", "PL"): "vieux", ("f", "PL"): "vieilles"},
"fou": {("m", "SG"): "fou", ("f", "SG"): "folle",
("m", "PL"): "fous", ("f", "PL"): "folles"},
"blanc": {("m", "SG"): "blanc", ("f", "SG"): "blanche",
("m", "PL"): "blancs", ("f", "PL"): "blanches"},
"long": {("m", "SG"): "long", ("f", "SG"): "longue",
("m", "PL"): "longs", ("f", "PL"): "longues"},
"bon": {("m", "SG"): "bon", ("f", "SG"): "bonne",
("m", "PL"): "bons", ("f", "PL"): "bonnes"},
}
def _rule_fem(a):
if a.endswith("e"):
return a
if a.endswith("er"):
return a[:-2] + "ère"
if a.endswith("eau"):
return a[:-3] + "elle"
if a.endswith("eux"):
return a[:-3] + "euse"
if a.endswith("f"):
return a[:-1] + "ve"
if a.endswith(("on", "en", "el", "eil", "et")):
return a + a[-1] + "e" # bon->bonne, ancien->ancienne, muet->muette
if a.endswith("c"):
return a[:-1] + "che" # blanc->blanche (public->publique via FIX)
return a + "e" # grand->grande, petit->petite, vert->verte
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
fix = _ADJ_FIX.get(lemma)
if fix and (g, num) in fix:
return fix[(g, num)], "lexicon"
d = _ADJS.get(lemma)
if d and d.get((g, num)):
return d[(g, num)], "lexicon"
# derive
msc = (d.get(("m", "SG")) if d else None) or lemma
if g == "m" and num == "SG":
return msc, "lexicon" if d else "rule"
fem = (d.get(("f", "SG")) if d else None) or _rule_fem(msc)
if g == "f" and num == "SG":
return fem, "lexicon" if (d and d.get(("f", "SG"))) else "rule"
if g == "m" and num == "PL":
if msc.endswith(("s", "x")):
return msc, "rule"
if msc.endswith("al"):
return msc[:-2] + "aux", "rule"
if msc.endswith("eau"):
return msc + "x", "rule"
return msc + "s", "rule"
# f|PL
return (fem if fem.endswith("s") else fem + "s"), "rule"
def lexicon_stats():
return {
"verb_source": "UniMorph French (github.com/unimorph/fra) + kaikki.org "
"irregulars (être + high-frequency)",
"noun_adj_source": "kaikki.org French (Wiktionary extract)",
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
"unimorph_verb_forms": len(_VERBS),
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
"irregular_verb_lemmas": len(_IRREGV),
"participle_lemmas": len(_PART),
"gerund_lemmas": len(_GER),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("parler", "ind", "present", "first", "singular", "parle"),
("être", "ind", "present", "third", "singular", "est"),
("avoir", "ind", "present", "first", "singular", "ai"),
("aller", "ind", "present", "third", "plural", "vont"),
("finir", "ind", "present", "first", "singular", "finis"),
("finir", "ind", "present", "first", "plural", "finissons"),
("manger", "ind", "present", "first", "plural", "mangeons"),
("faire", "ind", "future", "first", "singular", "ferai"),
("pouvoir", "sbjv", "present", "third", "singular", "puisse"),
("prendre", "ind", "passe_simple", "third", "singular", "prit"),
("vendre", "ind", "present", "third", "singular", "vend"),
("commencer", "ind", "imperfect", "first", "singular", "commençais"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
ok += got == exp
print(f" {flag}{lemma:10} {mood}/{tense:12} {per[:3]}.{num[:2]} -> {got:12} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" gender: maison=", noun_gender("maison"), "chat=", noun_gender("chat"),
"cheval=", noun_gender("cheval"), "nation=", noun_gender("nation"))
print(" plural: cheval->", inflect_noun("cheval", "plural"),
"| bateau->", inflect_noun("bateau", "plural"),
"| prix->", inflect_noun("prix", "plural"),
"| chat->", inflect_noun("chat", "plural"))
print(" adj: petit/f/sg->", inflect_adj("petit", "f", "singular"),
"| beau/f/sg->", inflect_adj("beau", "f", "singular"),
"| heureux/f/sg->", inflect_adj("heureux", "f", "singular"),
"| national/m/pl->", inflect_adj("national", "m", "plural"))
print(" part: aller/f/sg->", participle("aller", "f", "singular"),
"| prendre/f/pl->", participle("prendre", "f", "plural"),
"| finir/m/pl->", participle("finir", "m", "plural"))
print(" ger: manger->", gerund("manger"), "| finir->", gerund("finir"))
-588
View File
@@ -1,588 +0,0 @@
"""morphology_it_full.py — production-grade Italian morphological generator.
NOT a toy. Backed by three real, Wiktionary-lineage lexical sources:
VERBS
UniMorph Italian (github.com/unimorph/ita, CC-BY-SA 3.0)
10,009 verb lemmas × full paradigm, CLEAN orthography (no stress marks):
indicative present / imperfetto (PST;IPFV) / passato remoto (PST;PFV) /
futuro, condizionale (COND),
congiuntivo presente (SBJV;PRS) / imperfetto (SBJV;PST),
affirmative imperative, infinitive, gerundio (V.CVB;PRS),
past participle (masc-sg; fem/plural derived by vowel rule).
it_irreg_verbs.json 66 high-frequency verbs UniMorph MISSES
(essere, avere, potere, uscire, tenere, prendere, piacere, ), extracted
from kaikki.org Italian, filtered to standard forms, and DE-STRESSED to
real orthography (kaikki marks tonic stress everywhere: pàrlo->parlo,
avùto->avuto; final legit accents kept: sarò, è). Built by build_it_irreg.py.
This layer takes priority it supplies the two auxiliaries essere/avere,
which the whole passato-prossimo / essere-agreement system depends on.
NOUNS + ADJECTIVES kaikki.org Italian (Wiktionary extract, CC-BY-SA 3.0)
noun lemmas WITH inherent gender (head-template arg) + real (often irregular)
plural uomo->uomini, uovo->uova, dito->dita, città invariant resolved
PER LEMMA, never guessed.
adjective lemmas with real feminine + masc/fem plural (italiano->italiana/
italiani/italiane, felice->felici invariant).
Fallbacks (degrade, never crash, on OOV input):
verbs : rule generator for regular -are/-ere/-ire (with -care/-gare h-insertion
and -ciare/-giare/-iare i-drop spelling rules)
nouns : gender heuristic (endings) + rule pluralization (ambiguous -co/-go FLAGGED)
adjs : -o/-a/-e gender rule + rule pluralization
Confidence flag on every form:
"lexicon" from UniMorph / kaikki-irregular / kaikki noun-adj (trust: high)
"rule" deterministic rule (trust: medium)
"fallback" could not inflect; returned lemma / ambiguous (trust: low -> FLAG)
Public API (used by realizer_it.py):
conjugate(lemma, mood, tense, person, number) -> (form, conf)
participle(lemma, gender="m", number="singular") -> (form, conf)
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"
inflect_noun(lemma, number, gender=None) -> (form, conf)
inflect_adj(lemma, gender, number) -> (form, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "ita.unimorph")
_IRREG = os.path.join(_HERE, "data", "it_irreg_verbs.json")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_it.jsonl")
_CACHE = os.path.join(_HERE, "data", "it_morph_cache.pkl")
# ── (mood, tense) -> UniMorph feature set that must ALL be present ────────────────
_VERB_KEYMAP = {
("ind", "present"): {"IND", "PRS"},
("ind", "imperfect"): {"IND", "PST", "IPFV"},
("ind", "passato_remoto"): {"IND", "PST", "PFV"},
("ind", "future"): {"IND", "FUT"},
("ind", "conditional"): {"COND"},
("sbjv", "present"): {"SBJV", "PRS"},
("sbjv", "imperfect"): {"SBJV", "PST"},
("imp", "affirmative"): {"POS", "IMP"},
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── build verb lexicon from UniMorph ─────────────────────────────────────────────
def _build_verbs():
verbs = {} # (lemma, "mood|tense|person|number") -> form
part = {} # lemma -> masc-sg past participle
ger = {} # lemma -> gerundio
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f:
part.setdefault(lemma, form)
continue
if head == "V.CVB": # gerundio (converb, present)
if "PRS" in f:
ger.setdefault(lemma, form)
continue
if head != "V":
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
for (mood, tense), req in _VERB_KEYMAP.items():
# exact-set discipline: PST;PFV must not match PST;IPFV, etc.
if not req <= f:
continue
# guard IND;PST ambiguity: require the specific aspect feature
if tense == "imperfect" and "PFV" in f:
continue
if tense == "passato_remoto" and "IPFV" in f:
continue
# COND must not also be a subjunctive/imperative slot
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), form)
break
return verbs, part, ger
# ── kaikki nouns + adjectives ────────────────────────────────────────────────────
_EXCL_FORM_TAGS = {"alternative", "archaic", "obsolete", "dialectal", "regional",
"diminutive", "augmentative", "pejorative", "comparative",
"superlative", "misspelling", "rare", "informal", "literary",
"poetic", "error-unrecognized-form", "apocopic", "obsolete",
"construed", "collective"}
def _kaikki_gender(arg):
if not arg:
return None
a = str(arg).lower()
if a.startswith("f"):
return "f"
if a.startswith("m"):
return "m"
return None
def _build_nouns_adjs():
nouns = {} # lemma -> {"g","SG","PL"}
adjs = {} # lemma -> {("m","SG"),("f","SG"),("m","PL"),("f","PL")}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
pos = d.get("pos")
word = d.get("word", "")
if not word or " " in word:
continue
forms = d.get("forms", []) or []
if pos == "noun":
ht = d.get("head_templates") or []
g = None
if ht:
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
if g is None:
tags = d.get("tags") or []
if "feminine" in tags:
g = "f"
elif "masculine" in tags:
g = "m"
pl = None
for x in forms:
t = set(x.get("tags") or [])
if "plural" in t and not (t & _EXCL_FORM_TAGS):
fm = x.get("form")
if fm and " " not in fm and fm != "#":
pl = fm
break
if word not in nouns:
nouns[word] = {"g": g, "SG": word, "PL": pl}
else:
cur = nouns[word]
if cur.get("g") is None and g:
cur["g"] = g
if not cur.get("PL") and pl:
cur["PL"] = pl
elif pos == "adj":
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word)
for x in forms:
t = set(x.get("tags") or [])
fm = x.get("form")
if not fm or " " in fm or (t & _EXCL_FORM_TAGS):
continue
if "feminine" in t and "plural" in t:
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
elif "masculine" in t and "plural" in t:
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
elif "feminine" in t:
d0[("f", "SG")] = d0.get(("f", "SG")) or fm
elif "plural" in t: # invariant-gender adj (felice -> felici)
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
return nouns, adjs
def _build_cache():
verbs, part, ger = _build_verbs()
nouns, adjs = _build_nouns_adjs()
with open(_IRREG, encoding="utf-8") as fh:
irreg = json.load(fh)
data = {"verbs": verbs, "part": part, "ger": ger,
"nouns": nouns, "adjs": adjs, "irreg": irreg}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [_UNIMORPH, _KAIKKI, _IRREG]
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PART, _GER, _NOUNS, _ADJS, _IRREGV = (
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"],
_LEX["irreg"])
# ── regular-ending rule fallback ─────────────────────────────────────────────────
def _vclass(lemma):
if lemma.endswith("are"):
return "are"
if lemma.endswith("ere"):
return "ere"
if lemma.endswith("ire"):
return "ire"
return None
# endings [1sg,2sg,3sg,1pl,2pl,3pl]
_REG = {
("ind", "present", "are"): ["o", "i", "a", "iamo", "ate", "ano"],
("ind", "present", "ere"): ["o", "i", "e", "iamo", "ete", "ono"],
("ind", "present", "ire"): ["o", "i", "e", "iamo", "ite", "ono"],
("ind", "imperfect", "are"): ["avo", "avi", "ava", "avamo", "avate", "avano"],
("ind", "imperfect", "ere"): ["evo", "evi", "eva", "evamo", "evate", "evano"],
("ind", "imperfect", "ire"): ["ivo", "ivi", "iva", "ivamo", "ivate", "ivano"],
("ind", "passato_remoto", "are"): ["ai", "asti", "ò", "ammo", "aste", "arono"],
("ind", "passato_remoto", "ere"): ["ei", "esti", "é", "emmo", "este", "erono"],
("ind", "passato_remoto", "ire"): ["ii", "isti", "ì", "immo", "iste", "irono"],
("sbjv", "present", "are"): ["i", "i", "i", "iamo", "iate", "ino"],
("sbjv", "present", "ere"): ["a", "a", "a", "iamo", "iate", "ano"],
("sbjv", "present", "ire"): ["a", "a", "a", "iamo", "iate", "ano"],
("sbjv", "imperfect", "are"): ["assi", "assi", "asse", "assimo", "aste", "assero"],
("sbjv", "imperfect", "ere"): ["essi", "essi", "esse", "essimo", "este", "essero"],
("sbjv", "imperfect", "ire"): ["issi", "issi", "isse", "issimo", "iste", "issero"],
# imperative: 2sg,3sg(Lei),1pl,2pl,3pl (1sg has none)
("imp", "affirmative", "are"): [None, "a", "i", "iamo", "ate", "ino"],
("imp", "affirmative", "ere"): [None, "i", "a", "iamo", "ete", "ano"],
("imp", "affirmative", "ire"): [None, "i", "a", "iamo", "ite", "ano"],
}
# future / conditional attach to a stem = infinitive minus final -e, with
# -are -> -er (parlare->parler-), -ere/-ire keep (credere->creder-, dormir-)
_FUT = ["ò", "ai", "à", "emo", "ete", "anno"]
_COND = ["ei", "esti", "ebbe", "emmo", "este", "ebbero"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _fut_stem(lemma, vc):
body = lemma[:-3] # drop are/ere/ire
if vc == "are":
return body + "er"
return body + vc[0] + "r" # ere->er? no: keep vowel: creder-, dormir-
# NOTE corrected below
def _apply_are_spelling(stem, ending):
"""-care/-gare insert h before front endings; -ciare/-giare/-sciare/-iare drop i."""
front = ending[:1] in ("i", "e")
if stem.endswith(("c", "g")) and front:
return stem + "h" + ending
if stem.endswith(("ci", "gi", "sci")) and ending[:1] == "i":
return stem[:-1] + ending # mangi+iamo -> mangiamo
if stem.endswith("i") and ending[:1] == "i":
return stem[:-1] + ending # studi+iamo -> studiamo
return stem + ending
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
body = lemma[:-3]
i = _slot_idx(person, number)
if mood == "ind" and tense in ("future", "conditional"):
stem = body + "er" if vc == "are" else body + vc[0] + "r"
# ere: creder-, ire: dormir- -> body + 'e'/'i' + 'r'
if vc == "ere":
stem = body + "er"
elif vc == "ire":
stem = body + "ir"
end = (_FUT if tense == "future" else _COND)[i]
# spelling: -care/-gare -> cherò/gherò ; -ciare/-giare -> cerò/gerò
if vc == "are":
if body.endswith(("c", "g")):
stem = body + "her"
elif body.endswith(("ci", "gi", "sci")):
stem = body[:-1] + "er"
elif body.endswith("i"):
stem = body[:-1] + "er"
return stem + end
table = _REG.get((mood, tense, vc))
if not table:
return None
end = table[i]
if end is None:
return None
if vc == "are":
return _apply_are_spelling(body, end)
# -ere/-ire: guard against double-i (dormi+iamo -> dormiamo)
if body.endswith("i") and end[:1] == "i":
return body[:-1] + end
return body + end
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
"""Return (surface, confidence). mood in ind|sbjv|imp; tense per _VERB_KEYMAP."""
lemma = lemma.strip().lower()
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{number}"
ir = _IRREGV.get(lemma)
if ir and key in ir:
return ir[key], "lexicon"
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
if form:
return form, "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r:
return r, "rule"
return lemma, "fallback"
# ── PUBLIC: participle + gerund ──────────────────────────────────────────────────
def _participle_msg(lemma):
"""Return (masc-sg participle, source) or (None, None)."""
ir = _IRREGV.get(lemma)
if ir and "part" in ir:
return ir["part"], "lexicon"
if lemma in _PART:
return _PART[lemma], "lexicon"
return None, None
def participle(lemma, gender="m", number="singular"):
"""Past participle with gender/number agreement (for essere-perfect & passives).
UniMorph/irregular give masc-sg; fem/plural derived by final-vowel swap
(-o -> -a/-i/-e), valid for regular -ato/-uto/-ito AND irregulars
(preso->presa/presi/prese, aperto->aperta/aperti/aperte, morto->morta/...)."""
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
msg, src = _participle_msg(lemma)
conf = "lexicon"
if msg is None:
vc = _vclass(lemma)
if vc == "are":
msg = lemma[:-3] + "ato"
elif vc == "ere":
msg = lemma[:-3] + "uto"
elif vc == "ire":
msg = lemma[:-3] + "ito"
else:
return lemma, "fallback"
conf = "rule"
# agreement: only -o participles inflect for gender+number
if msg.endswith("o"):
stem = msg[:-1]
suf = {"m|SG": "o", "f|SG": "a", "m|PL": "i", "f|PL": "e"}[f"{g}|{num}"]
return stem + suf, conf
return msg, conf # non -o participle: leave as-is (rare)
def gerund(lemma):
lemma = lemma.strip().lower()
ir = _IRREGV.get(lemma)
if ir and "ger" in ir:
return ir["ger"], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
vc = _vclass(lemma)
if vc == "are":
return lemma[:-3] + "ando", "rule"
if vc in ("ere", "ire"):
return lemma[:-3] + "endo", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
_FEM_SUF = ("zione", "sione", "gione", "", "", "trice", "aggine", "udine",
"igine", "ie", "essa", "izia", "ezza")
_MASC_SUF = ("ore", "ame", "iere", "ale", "ile")
def _gender_heuristic(noun):
for suf in _FEM_SUF:
if noun.endswith(suf):
return "f"
for suf in _MASC_SUF:
if noun.endswith(suf):
return "m"
if noun.endswith("o"):
return "m"
if noun.endswith("a"):
return "f"
if noun.endswith("à") or noun.endswith("ù"):
return "f"
return "m" # -e and consonant-final loanwords default masculine
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g") in ("m", "f"):
return d["g"]
return _gender_heuristic(lemma)
def _rule_plural(noun, gender):
"""Deterministic Italian pluralization. Returns (form, ok); ok=False FLAGS an
ambiguous case the lexicon would normally resolve (-co/-go palatalization)."""
if not noun:
return noun, True
# invariant: accented final vowel, consonant-final, monosyllable, -i final
if noun[-1:] in ("à", "è", "é", "ì", "í", "ò", "ó", "ù", "ú"):
return noun, True
if noun[-1:] not in ("a", "e", "o", "i", "u"):
return noun, True # consonant-final loanword: invariant
if noun.endswith("i"):
return noun, True # e.g. crisi, analisi: invariant
if noun.endswith("io"):
return noun[:-2] + "i", True # figlio->figli (unstressed i)
if noun.endswith("cia") or noun.endswith("gia"):
# vowel before cia/gia -> -cie/-gie ; consonant -> -ce/-ge (approx)
return noun[:-2] + "e", True # arancia->arance (majority)
if noun.endswith("ca"):
return noun[:-2] + "che", True # amica->amiche
if noun.endswith("ga"):
return noun[:-2] + "ghe", True
if noun.endswith("co"):
return noun[:-2] + "chi", False # AMBIGUOUS (amico->amici) -> flag
if noun.endswith("go"):
return noun[:-2] + "ghi", False # AMBIGUOUS (psicologo->psicologi)
if noun.endswith("a"):
return noun[:-1] + "e", True # casa->case (m -a: -i, but rare)
if noun.endswith("o"):
return noun[:-1] + "i", True # libro->libri
if noun.endswith("e"):
return noun[:-1] + "i", True # cane->cani, chiave->chiavi
return noun, True
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if number == "singular":
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
if d and d.get("PL"):
return d["PL"], "lexicon"
g = gender or noun_gender(lemma)
form, ok = _rule_plural(lemma, g)
return form, ("rule" if ok else "fallback")
# adjectives whose kaikki entries are unreliable (messy inflection templates):
# supply audited regular agreement forms (prenominal apocope handled in realizer).
_ADJ_FIX = {
"bello": {("m", "SG"): "bello", ("f", "SG"): "bella",
("m", "PL"): "belli", ("f", "PL"): "belle"},
"quello": {("m", "SG"): "quello", ("f", "SG"): "quella",
("m", "PL"): "quelli", ("f", "PL"): "quelle"},
}
# ── PUBLIC: adjective agreement ──────────────────────────────────────────────────
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
fix = _ADJ_FIX.get(lemma)
if fix and (g, num) in fix:
return fix[(g, num)], "lexicon"
d = _ADJS.get(lemma)
if d:
form = d.get((g, num))
if form:
return form, "lexicon"
sg = d.get((g, "SG")) or d.get(("m", "SG")) or lemma
if num == "PL":
pl, ok = _rule_plural(sg, g)
return pl, ("rule" if ok else "fallback")
return sg, "lexicon"
# rule fallback
a = lemma
if a.endswith("o"): # -o/-a/-i/-e class
base = a[:-1]
suf = {"m|SG": "o", "f|SG": "a", "m|PL": "i", "f|PL": "e"}[f"{g}|{num}"]
return base + suf, "rule"
if a.endswith("e"): # felice-class: SG invariant, PL -i
if num == "PL":
return a[:-1] + "i", "rule"
return a, "rule"
if num == "PL":
p, ok = _rule_plural(a, g)
return p, ("rule" if ok else "fallback")
return a, "rule"
def lexicon_stats():
return {
"verb_source": "UniMorph Italian (github.com/unimorph/ita) + kaikki.org "
"irregulars (de-stressed)",
"noun_adj_source": "kaikki.org Italian (Wiktionary extract)",
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
"unimorph_verb_forms": len(_VERBS),
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
"irregular_verb_lemmas": len(_IRREGV),
"participle_lemmas": len(_PART),
"gerund_lemmas": len(_GER),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("parlare", "ind", "present", "first", "singular", "parlo"),
("essere", "ind", "present", "third", "singular", "è"),
("avere", "ind", "present", "first", "singular", "ho"),
("mangiare", "ind", "present", "second", "singular", "mangi"),
("finire", "ind", "present", "first", "singular", "finisco"),
("andare", "ind", "present", "third", "plural", "vanno"),
("fare", "ind", "future", "first", "singular", "farò"),
("potere", "sbjv", "present", "third", "singular", "possa"),
("prendere", "ind", "passato_remoto", "first", "singular", "presi"),
("cercare", "ind", "present", "second", "singular", "cerchi"),
("dormire", "ind", "present", "third", "plural", "dormono"),
("credere", "ind", "future", "first", "singular", "crederò"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
ok += got == exp
print(f" {flag}{lemma:9} {mood}/{tense:14} {per[:3]}.{num[:2]} -> {got:12} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" gender: casa=", noun_gender("casa"), "problema=", noun_gender("problema"),
"mano=", noun_gender("mano"), "città=", noun_gender("città"),
"cane=", noun_gender("cane"))
print(" plural: uomo->", inflect_noun("uomo", "plural"),
"| uovo->", inflect_noun("uovo", "plural"),
"| città->", inflect_noun("città", "plural"),
"| amico->", inflect_noun("amico", "plural"),
"| casa->", inflect_noun("casa", "plural"))
print(" adj: italiano/f/pl->", inflect_adj("italiano", "f", "plural"),
"| felice/m/pl->", inflect_adj("felice", "m", "plural"),
"| bello/f/sg->", inflect_adj("bello", "f", "singular"))
print(" part: aprire/f/sg->", participle("aprire", "f", "singular"),
"| prendere/m/pl->", participle("prendere", "m", "plural"),
"| andare/f/sg->", participle("andare", "f", "singular"))
print(" ger: fare->", gerund("fare"), "| parlare->", gerund("parlare"))
-666
View File
@@ -1,666 +0,0 @@
# -*- coding: utf-8 -*-
"""morphology_lat_full.py — production-grade Latin morphological generator.
Latin is the FLAGSHIP dead-language realizer. It rides the *architecture* of the
Romance/Italic engine (the same Realization / spec-driven design and the UniMorph
loader pattern from morphology_it_full.py) but with the CASE SYSTEM RESTORED
the feature Romance lost. Latin therefore exercises machinery the modern Romance
siblings never needed: 5 declensions x 6 cases x 2 numbers x 3 genders, plus a
4-conjugation verb system with tense/mood/voice.
DATA (real, attested no fabrication):
NOUNS + ADJECTIVES UniMorph Latin (github.com/unimorph/lat, CC-BY-SA 3.0)
163,182 N forms across ~thousands of lemmas, each with the full case paradigm
N;NOM/GEN/DAT/ACC/ABL/VOC;SG/PL (real inflected forms, WITH macrons:
puella->puellam, rēx->rēgis, corpus->corporis).
244,197 ADJ forms with case x GENDER x number, incl. UniMorph's combined
tags (GEN+DAT, MASC+FEM, MASC+FEM+NEUT) which are split on load.
462,668 V.PTCP forms (participles) also carry case/gender/number.
UniMorph N tags DO NOT encode inherent gender, so noun gender is inferred
from the declension (nom-sg + gen-sg endings) with a curated exceptions
map the standard, attestable rule (1st decl -a/-ae = fem, 2nd -us/-i =
masc, -um = neut, ...).
VERBS RULE ENGINE (honest gap: UniMorph Latin's verb list is a 947-lemma
sample of rare/prefixed verbs that MISSES every core textbook verb amō,
videō, sum, regō, ... are all absent). Latin conjugation is, however, highly
regular, so verbs are generated by a deterministic 4-conjugation engine over
curated principal parts (present / perfect / supine stems), sourced from
standard references. Irregulars (sum, possum, , ferō, volō, nōlō, mālō)
are curated full tables. Forms are flagged "rule" (not "lexicon") for honesty.
Confidence flag on every form (same contract as the Romance engine):
"lexicon" from UniMorph (trust: high)
"rule" deterministic morphology rule (trust: medium)
"fallback" could not inflect; returned lemma (trust: low -> FLAG)
Public API (used by realizer_lat.py):
decline_noun(lemma, case, number) -> (form, conf)
noun_gender(lemma) -> "m"|"f"|"n"
decline_adj(lemma, case, gender, number) -> (form, conf)
conjugate(lemma, tense, mood, voice, person, number) -> (form, conf)
participle(lemma, kind, case, gender, number) -> (form, conf) # kind: prs|pfv|fut
infinitive(lemma, tense="present", voice="active") -> (form, conf)
lexicon_stats() -> dict
"""
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "lat.unimorph")
_CACHE = os.path.join(_HERE, "data", "lat_morph_cache.pkl")
_CASES = ("NOM", "GEN", "DAT", "ACC", "ABL", "VOC")
_CASE_MAP = {"nom": "NOM", "gen": "GEN", "dat": "DAT", "acc": "ACC",
"abl": "ABL", "voc": "VOC"}
_NUM = {"singular": "SG", "plural": "PL"}
_GEN = {"m": "MASC", "f": "FEM", "n": "NEUT"}
# ── UniMorph loader: noun + adjective + participle case paradigms ────────────────
def _build_cache():
nouns = {} # lemma -> {(CASE, NUM): form}
adjs = {} # lemma -> {(CASE, GEN, NUM): form}
ptcps = {} # lemma -> {(CASE, GEN, NUM): form} (from V.PTCP; keyed loosely)
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
feats = tag.split(";")
head = feats[0]
fs = set(feats)
case = next((c for c in _CASES if c in fs), None)
# handle combined case tags like GEN+DAT
if case is None:
for f in feats:
if "+" in f and any(c in f.split("+") for c in _CASES):
case = [c for c in _CASES if c in f.split("+")]
break
num = "SG" if "SG" in fs else ("PL" if "PL" in fs else None)
if case is None or num is None:
continue
cases = case if isinstance(case, list) else [case]
if head == "N":
d = nouns.setdefault(lemma, {})
for c in cases:
d.setdefault((c, num), form)
elif head == "ADJ":
# gender may be combined: MASC+FEM+NEUT, MASC+FEM
genders = []
for g in ("MASC", "FEM", "NEUT"):
if any(g == x or (g in x.split("+")) for x in feats):
genders.append(g)
if not genders:
genders = ["MASC", "FEM", "NEUT"]
d = adjs.setdefault(lemma, {})
for c in cases:
for g in genders:
d.setdefault((c, g, num), form)
data = {"nouns": nouns, "adjs": adjs, "ptcps": ptcps}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE) and os.path.exists(_UNIMORPH):
if os.path.getmtime(_CACHE) >= os.path.getmtime(_UNIMORPH):
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_NOUNS, _ADJS = _LEX["nouns"], _LEX["adjs"]
# ── noun gender inference (declension-based, curated exceptions) ─────────────────
# Real, attestable rule: gender follows declension + nominative shape, with the
# standard closed set of exceptions.
_GENDER_EXC = {
# 1st-declension masculines (people/agents)
"agricola": "m", "poēta": "m", "nauta": "m", "incola": "m", "scrība": "m",
"auriga": "m", "pīrāta": "m", "athlēta": "m",
# 2nd-declension neuters / feminines
"vīrus": "n", "vulgus": "n", "pelagus": "n", "humus": "f",
# common 3rd-declension whose gender the ending would mispredict
"rēx": "m", "dux": "m", "mīles": "m", "pater": "m", "frāter": "m",
"homō": "m", "leō": "m", "sōl": "m", "mōns": "m", "pōns": "m", "fōns": "m",
"sanguis": "m", "ōrdō": "m", "sermō": "m", "amor": "m", "dolor": "m",
"labor": "m", "timor": "m", "honor": "m", "color": "m", "pēs": "m",
"dēns": "m", "flōs": "m", "mōs": "m", "mensis": "m", "orbis": "m",
"piscis": "m", "ignis": "m", "collis": "m", "grex": "m", "prīnceps": "m",
"māter": "f", "soror": "f", "uxor": "f", "mulier": "f", "virgō": "f",
"urbs": "f", "arx": "f", "pāx": "f", "lēx": "f", "lūx": "f", "vōx": "f",
"nox": "f", "nix": "f", "vīs": "f", "salūs": "f", "virtūs": "f",
"aetās": "f", "cīvitās": "f", "lībertās": "f", "vēritās": "f", "voluptās": "f",
"nātiō": "f", "ratiō": "f", "ōrātiō": "f", "legiō": "f", "regiō": "f",
"mens": "f", "gens": "f", "ars": "f", "pars": "f", "mors": "f", "sors": "f",
"nāvis": "f", "turris": "f", "avis": "f", "vallis": "f", "classis": "f",
"corpus": "n", "tempus": "n", "opus": "n", "genus": "n", "onus": "n",
"pectus": "n", "latus": "n", "vulnus": "n", "scelus": "n", "sīdus": "n",
"caput": "n", "iter": "n", "flūmen": "n", "nōmen": "n", "carmen": "n",
"agmen": "n", "certāmen": "n", "lūmen": "n", "ōmen": "n", "cōgnōmen": "n",
"mare": "n", "animal": "n", "exemplar": "n", "rēte": "n",
# 4th-declension exceptions
"manus": "f", "domus": "f", "tribus": "f", "porticus": "f", "īdūs": "f",
"cornū": "n", "genū": "n", "gelū": "n", "verū": "n",
# 5th-declension
"diēs": "m", "merīdiēs": "m",
}
def _infer_gender(lemma):
if lemma in _GENDER_EXC:
return _GENDER_EXC[lemma]
d = _NOUNS.get(lemma)
nom = d.get(("NOM", "SG")) if d else lemma
gen = d.get(("GEN", "SG")) if d else None
nom = nom or lemma
# 5th declension: gen -eī / -ēī
if gen and (gen.endswith("") or gen.endswith("ēī")):
return "f"
# 1st declension: nom -a, gen -ae
if nom.endswith("a") and (not gen or gen.endswith("ae")):
return "f"
# 2nd declension neuter: nom -um
if nom.endswith("um"):
return "n"
# 2nd declension masc: nom -us/-er/-ir, gen -ī
if (nom.endswith("us") or nom.endswith("er") or nom.endswith("ir")) and \
(not gen or gen.endswith("ī")):
return "m"
# 4th declension: gen -ūs
if gen and gen.endswith("ūs"):
return "n" if nom.endswith("ū") else "m"
# 3rd declension neuters by common nom endings
if nom.endswith(("men", "us", "ur", "al", "ar", "e", "ma")):
# -us here is 3rd-decl neuter type (corpus) only if gen shows -oris/-eris
if nom.endswith("us") and gen and (gen.endswith("oris") or gen.endswith("eris")
or gen.endswith("uris")):
return "n"
if nom.endswith(("men", "al", "ar", "e")):
return "n"
# default 3rd-declension: masculine (most common)
return "m"
_GENDER_CACHE = {}
def noun_gender(lemma):
lemma = lemma.strip()
if lemma not in _GENDER_CACHE:
_GENDER_CACHE[lemma] = _infer_gender(lemma)
return _GENDER_CACHE[lemma]
# ── PUBLIC: noun declension ─────────────────────────────────────────────────────
def decline_noun(lemma, case, number):
lemma = lemma.strip()
C = _CASE_MAP.get(case, case.upper())
N = _NUM.get(number, number)
d = _NOUNS.get(lemma)
if d and (C, N) in d:
return d[(C, N)], "lexicon"
# abl sg often == the -e/-o form; try nom fallback
if d:
# try VOC==NOM, ACC neuter==NOM etc are already in data; last resort lemma
return lemma, "fallback"
return lemma, "fallback"
# ── PUBLIC: adjective declension ────────────────────────────────────────────────
def decline_adj(lemma, case, gender, number):
lemma = lemma.strip()
C = _CASE_MAP.get(case, case.upper())
G = _GEN.get(gender, gender.upper())
N = _NUM.get(number, number)
d = _ADJS.get(lemma)
if d and (C, G, N) in d:
return d[(C, G, N)], "lexicon"
# try other gender (some adjs listed only under MASC+FEM etc handled at load)
if d:
for altG in ("MASC", "FEM", "NEUT"):
if (C, altG, N) in d:
return d[(C, altG, N)], "lexicon"
return lemma, "fallback"
return lemma, "fallback"
# ═══════════════════════════════════════════════════════════════════════════════
# VERB RULE ENGINE (4 conjugations + curated irregulars)
# ═══════════════════════════════════════════════════════════════════════════════
# Curated principal parts for common attested verbs:
# lemma -> (conj, present_stem, perfect_stem, supine_stem)
# conj in {1,2,3,"3io",4}. Stems carry macrons (matching UniMorph orthography).
_VERBS = {
"amō": (1, "am", "amāv", "amāt"),
"laudō": (1, "laud", "laudāv", "laudāt"),
"portō": (1, "port", "portāv", "portāt"),
"vocō": (1, "voc", "vocāv", "vocāt"),
"": (1, "d", "ded", "dat"),
"spectō": (1, "spect", "spectāv", "spectāt"),
"pugnō": (1, "pugn", "pugnāv", "pugnāt"),
"labōrō": (1, "labōr", "labōrāv", "labōrāt"),
"necō": (1, "nec", "necāv", "necāt"),
"parō": (1, "par", "parāv", "parāt"),
"cōgitō": (1, "cōgit", "cōgitāv", "cōgitāt"),
"habitō": (1, "habit", "habitāv", "habitāt"),
"nārrō": (1, "nārr", "nārrāv", "nārrāt"),
"servō": (1, "serv", "servāv", "servāt"),
"superō": (1, "super", "superāv", "superāt"),
"oppugnō": (1, "oppugn", "oppugnāv", "oppugnāt"),
"ambulō": (1, "ambul", "ambulāv", "ambulāt"),
"clāmō": (1, "clām", "clāmāv", "clāmāt"),
"vulnerō": (1, "vulner", "vulnerāv", "vulnerāt"),
"aedificō": (1, "aedific", "aedificāv", "aedificāt"),
"expugnō": (1, "expugn", "expugnāv", "expugnāt"),
"dēfendō": (3, "dēfend", "dēfend", "dēfēns"),
"petō": (3, "pet", "petīv", "petīt"),
"occīdō": (3, "occīd", "occīd", "occīs"),
"interficiō": ("3io", "interfic", "interfēc", "interfect"),
"timeō": (2, "tim", "timu", None),
"iaceō": (2, "iac", "iacu", None),
"pāreō": (2, "pār", "pāru", "pārit"),
"respondeō": (2, "respond", "respond", "respōns"),
"vertō": (3, "vert", "vert", "vers"),
"ostendō": (3, "ostend", "ostend", "ostent"),
"cōnstituō": (3, "cōnstitu", "cōnstitu", "cōnstitūt"),
"cōgnōscō": (3, "cōgnōsc", "cōgnōv", "cōgnit"),
"crēdō": (3, "crēd", "crēdid", "crēdit"),
"ēdūcō": (3, "ēdūc", "ēdūx", "ēduct"),
"cōnservō": (1, "cōnserv", "cōnservāv", "cōnservāt"),
"iuvō": (1, "iuv", "iūv", "iūt"),
"dēbeō": (2, "dēb", "dēbu", "dēbit"),
"moneō": (2, "mon", "monu", "monit"),
"videō": (2, "vid", "vīd", "vīs"),
"habeō": (2, "hab", "habu", "habit"),
"teneō": (2, "ten", "tenu", "tent"),
"timeō": (2, "tim", "timu", None),
"terreō": (2, "terr", "terru", "territ"),
"dēleō": (2, "dēl", "dēlēv", "dēlēt"),
"iubeō": (2, "iub", "iuss", "iuss"),
"maneō": (2, "man", "māns", "māns"),
"moveō": (2, "mov", "mōv", "mōt"),
"doceō": (2, "doc", "docu", "doct"),
"sedeō": (2, "sed", "sēd", "sess"),
"rīdeō": (2, "rīd", "rīs", "rīs"),
"regō": (3, "reg", "rēx", "rēct"),
"dūcō": (3, "dūc", "dūx", "duct"),
"scrībō": (3, "scrīb", "scrīps", "scrīpt"),
"mittō": (3, "mitt", "mīs", "miss"),
"pōnō": (3, "pōn", "posu", "posit"),
"agō": (3, "ag", "ēg", "āct"),
"dīcō": (3, "dīc", "dīx", "dict"),
"gerō": (3, "ger", "gess", "gest"),
"vincō": (3, "vinc", "vīc", "vict"),
"petō": (3, "pet", "petīv", "petīt"),
"legō": (3, "leg", "lēg", "lēct"),
"currō": (3, "curr", "cucurr", "curs"),
"vīvō": (3, "vīv", "vīx", "vīct"),
"quaerō": (3, "quaer", "quaesīv", "quaesīt"),
"trahō": (3, "trah", "trāx", "tract"),
"claudō": (3, "claud", "claus", "claus"),
"cōgō": (3, "cōg", "coēg", "coāct"),
"relinquō": (3, "relinqu", "relīqu", "relict"),
"capiō": ("3io", "cap", "cēp", "capt"),
"faciō": ("3io", "fac", "fēc", "fact"),
"iaciō": ("3io", "iac", "iēc", "iact"),
"rapiō": ("3io", "rap", "rapu", "rapt"),
"fugiō": ("3io", "fug", "fūg", "fugit"),
"cupiō": ("3io", "cup", "cupīv", "cupīt"),
"accipiō": ("3io", "accip", "accēp", "accept"),
"audiō": (4, "aud", "audīv", "audīt"),
"veniō": (4, "ven", "vēn", "vent"),
"sciō": (4, "sc", "scīv", "scīt"),
"sentiō": (4, "sent", "sēns", "sēns"),
"mūniō": (4, "mūn", "mūnīv", "mūnīt"),
"dormiō": (4, "dorm", "dormīv", "dormīt"),
"aperiō": (4, "aper", "aperu", "apert"),
"inveniō": (4, "inven", "invēn", "invent"),
}
# ── Present-system paradigms: full ending tables per conjugation, attached to the
# bare present stem (pstem). Hardcoded from the standard grammar with correct
# macrons/vowel-lengths — deterministic and independently verifiable. Keys:
# (tense, mood, voice) -> {conj: [1sg,2sg,3sg,1pl,2pl,3pl]}
_PARADIGM = {
("present", "ind", "active"): {
1: ["ō", "ās", "at", "āmus", "ātis", "ant"],
2: ["", "ēs", "et", "ēmus", "ētis", "ent"],
3: ["ō", "is", "it", "imus", "itis", "unt"],
"3io": ["", "is", "it", "imus", "itis", "iunt"],
4: ["", "īs", "it", "īmus", "ītis", "iunt"],
},
("present", "ind", "passive"): {
1: ["or", "āris", "ātur", "āmur", "āminī", "antur"],
2: ["eor", "ēris", "ētur", "ēmur", "ēminī", "entur"],
3: ["or", "eris", "itur", "imur", "iminī", "untur"],
"3io": ["ior", "eris", "itur", "imur", "iminī", "iuntur"],
4: ["ior", "īris", "ītur", "īmur", "īminī", "iuntur"],
},
("imperfect", "ind", "active"): {
1: ["ābam", "ābās", "ābat", "ābāmus", "ābātis", "ābant"],
2: ["ēbam", "ēbās", "ēbat", "ēbāmus", "ēbātis", "ēbant"],
3: ["ēbam", "ēbās", "ēbat", "ēbāmus", "ēbātis", "ēbant"],
"3io": ["iēbam", "iēbās", "iēbat", "iēbāmus", "iēbātis", "iēbant"],
4: ["iēbam", "iēbās", "iēbat", "iēbāmus", "iēbātis", "iēbant"],
},
("imperfect", "ind", "passive"): {
1: ["ābar", "ābāris", "ābātur", "ābāmur", "ābāminī", "ābantur"],
2: ["ēbar", "ēbāris", "ēbātur", "ēbāmur", "ēbāminī", "ēbantur"],
3: ["ēbar", "ēbāris", "ēbātur", "ēbāmur", "ēbāminī", "ēbantur"],
"3io": ["iēbar", "iēbāris", "iēbātur", "iēbāmur", "iēbāminī", "iēbantur"],
4: ["iēbar", "iēbāris", "iēbātur", "iēbāmur", "iēbāminī", "iēbantur"],
},
("future", "ind", "active"): {
1: ["ābō", "ābis", "ābit", "ābimus", "ābitis", "ābunt"],
2: ["ēbō", "ēbis", "ēbit", "ēbimus", "ēbitis", "ēbunt"],
3: ["am", "ēs", "et", "ēmus", "ētis", "ent"],
"3io": ["iam", "iēs", "iet", "iēmus", "iētis", "ient"],
4: ["iam", "iēs", "iet", "iēmus", "iētis", "ient"],
},
("future", "ind", "passive"): {
1: ["ābor", "āberis", "ābitur", "ābimur", "ābiminī", "ābuntur"],
2: ["ēbor", "ēberis", "ēbitur", "ēbimur", "ēbiminī", "ēbuntur"],
3: ["ar", "ēris", "ētur", "ēmur", "ēminī", "entur"],
"3io": ["iar", "iēris", "iētur", "iēmur", "iēminī", "ientur"],
4: ["iar", "iēris", "iētur", "iēmur", "iēminī", "ientur"],
},
("present", "sbjv", "active"): {
1: ["em", "ēs", "et", "ēmus", "ētis", "ent"],
2: ["eam", "eās", "eat", "eāmus", "eātis", "eant"],
3: ["am", "ās", "at", "āmus", "ātis", "ant"],
"3io": ["iam", "iās", "iat", "iāmus", "iātis", "iant"],
4: ["iam", "iās", "iat", "iāmus", "iātis", "iant"],
},
("present", "sbjv", "passive"): {
1: ["er", "ēris", "ētur", "ēmur", "ēminī", "entur"],
2: ["ear", "eāris", "eātur", "eāmur", "eāminī", "eantur"],
3: ["ar", "āris", "ātur", "āmur", "āminī", "antur"],
"3io": ["iar", "iāris", "iātur", "iāmur", "iāminī", "iantur"],
4: ["iar", "iāris", "iātur", "iāmur", "iāminī", "iantur"],
},
("imperfect", "sbjv", "active"): {
1: ["ārem", "ārēs", "āret", "ārēmus", "ārētis", "ārent"],
2: ["ērem", "ērēs", "ēret", "ērēmus", "ērētis", "ērent"],
3: ["erem", "erēs", "eret", "erēmus", "erētis", "erent"],
"3io": ["erem", "erēs", "eret", "erēmus", "erētis", "erent"],
4: ["īrem", "īrēs", "īret", "īrēmus", "īrētis", "īrent"],
},
("imperfect", "sbjv", "passive"): {
1: ["ārer", "ārēris", "ārētur", "ārēmur", "ārēminī", "ārentur"],
2: ["ērer", "ērēris", "ērētur", "ērēmur", "ērēminī", "ērentur"],
3: ["erer", "erēris", "erētur", "erēmur", "erēminī", "erentur"],
"3io": ["erer", "erēris", "erētur", "erēmur", "erēminī", "erentur"],
4: ["īrer", "īrēris", "īrētur", "īrēmur", "īrēminī", "īrentur"],
},
}
# perfect-active endings (added to perfect stem) — same for all conjugations
_PERF_ACT = {
("perfect", "ind"): ["ī", "istī", "it", "imus", "istis", "ērunt"],
("pluperfect", "ind"): ["eram", "erās", "erat", "erāmus", "erātis", "erant"],
("futureperfect", "ind"): ["erō", "eris", "erit", "erimus", "eritis", "erint"],
("perfect", "sbjv"): ["erim", "erīs", "erit", "erīmus", "erītis", "erint"],
("pluperfect", "sbjv"):["issem", "issēs", "isset", "issēmus", "issētis", "issent"],
}
def _idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _present_system(conj, pstem, tense, mood, voice, person, number):
"""Generate a present-system form (present/imperfect/future ind & subj)."""
table = _PARADIGM.get((tense, mood, voice))
if not table or conj not in table:
return None
return pstem + table[conj][_idx(person, number)]
def _active_infinitive_stem(conj, pstem):
return {1: pstem + "ā", 2: pstem + "ē", 3: pstem + "e",
"3io": pstem + "e", 4: pstem + "ī"}[conj]
_IRREG = {
"sum": {
("present", "ind", "active"): ["sum", "es", "est", "sumus", "estis", "sunt"],
("imperfect", "ind", "active"): ["eram", "erās", "erat", "erāmus", "erātis", "erant"],
("future", "ind", "active"): ["erō", "eris", "erit", "erimus", "eritis", "erunt"],
("perfect", "ind", "active"): ["fuī", "fuistī", "fuit", "fuimus", "fuistis", "fuērunt"],
("pluperfect", "ind", "active"): ["fueram", "fuerās", "fuerat", "fuerāmus", "fuerātis", "fuerant"],
("present", "sbjv", "active"): ["sim", "sīs", "sit", "sīmus", "sītis", "sint"],
("imperfect", "sbjv", "active"): ["essem", "essēs", "esset", "essēmus", "essētis", "essent"],
},
"possum": {
("present", "ind", "active"): ["possum", "potes", "potest", "possumus", "potestis", "possunt"],
("imperfect", "ind", "active"): ["poteram", "poterās", "poterat", "poterāmus", "poterātis", "poterant"],
("future", "ind", "active"): ["poterō", "poteris", "poterit", "poterimus", "poteritis", "poterunt"],
("perfect", "ind", "active"): ["potuī", "potuistī", "potuit", "potuimus", "potuistis", "potuērunt"],
("present", "sbjv", "active"): ["possim", "possīs", "possit", "possīmus", "possītis", "possint"],
},
"": {
("present", "ind", "active"): ["", "īs", "it", "īmus", "ītis", "eunt"],
("imperfect", "ind", "active"): ["ībam", "ībās", "ībat", "ībāmus", "ībātis", "ībant"],
("future", "ind", "active"): ["ībō", "ībis", "ībit", "ībimus", "ībitis", "ībunt"],
("perfect", "ind", "active"): ["", "īstī", "iit", "iimus", "īstis", "iērunt"],
("present", "sbjv", "active"): ["eam", "eās", "eat", "eāmus", "eātis", "eant"],
},
"volō": {
("present", "ind", "active"): ["volō", "vīs", "vult", "volumus", "vultis", "volunt"],
("imperfect", "ind", "active"): ["volēbam", "volēbās", "volēbat", "volēbāmus", "volēbātis", "volēbant"],
("future", "ind", "active"): ["volam", "volēs", "volet", "volēmus", "volētis", "volent"],
("perfect", "ind", "active"): ["voluī", "voluistī", "voluit", "voluimus", "voluistis", "voluērunt"],
("present", "sbjv", "active"): ["velim", "velīs", "velit", "velīmus", "velītis", "velint"],
},
"nōlō": {
("present", "ind", "active"): ["nōlō", "nōn vīs", "nōn vult", "nōlumus", "nōn vultis", "nōlunt"],
("present", "sbjv", "active"): ["nōlim", "nōlīs", "nōlit", "nōlīmus", "nōlītis", "nōlint"],
},
"ferō": {
("present", "ind", "active"): ["ferō", "fers", "fert", "ferimus", "fertis", "ferunt"],
("imperfect", "ind", "active"): ["ferēbam", "ferēbās", "ferēbat", "ferēbāmus", "ferēbātis", "ferēbant"],
("future", "ind", "active"): ["feram", "ferēs", "feret", "ferēmus", "ferētis", "ferent"],
("perfect", "ind", "active"): ["tulī", "tulistī", "tulit", "tulimus", "tulistis", "tulērunt"],
("present", "sbjv", "active"): ["feram", "ferās", "ferat", "ferāmus", "ferātis", "ferant"],
},
}
def conjugate(lemma, tense, mood, voice="active", person="third", number="singular"):
"""Return (surface, confidence). Perfect-passive forms are periphrastic and
handled in the realizer (sum + PPP); this returns synthetic forms only."""
lemma = lemma.strip()
i = _idx(person, number)
ir = _IRREG.get(lemma)
if ir:
tbl = ir.get((tense, mood, voice)) or ir.get((tense, mood, "active"))
if tbl and tbl[i]:
return tbl[i], "rule"
v = _VERBS.get(lemma)
if not v:
v = _infer_principal_parts(lemma)
if not v:
return lemma, "fallback"
conj, pstem, perfstem, supstem = v
# imperative (present active) 2sg / 2pl
if mood == "imp":
return _imperative(conj, pstem, person, number), "rule"
# perfect-system active
if tense in ("perfect", "pluperfect", "futureperfect") and voice == "active":
if not perfstem:
return lemma, "fallback"
end = _PERF_ACT.get((tense, mood))
if end:
return perfstem + end[i], "rule"
# present-system (active + passive)
if tense in ("present", "imperfect", "future"):
form = _present_system(conj, pstem, tense, mood, voice, person, number)
if form:
return form, "rule"
return lemma, "fallback"
def _imperative(conj, pstem, person, number):
if number == "singular":
return {1: pstem + "ā", 2: pstem + "ē", 3: pstem + "e",
"3io": pstem + "e", 4: pstem + "ī"}[conj]
return {1: pstem + "āte", 2: pstem + "ēte", 3: pstem + "ite",
"3io": pstem + "ite", 4: pstem + "īte"}[conj]
def _infer_principal_parts(lemma):
"""OOV fallback: infer conjugation + stems from the 1sg-present citation form.
Perfect/supine stems are guessed regularly (often wrong for 3rd conj) and the
resulting forms are still returned as 'rule' but the realizer down-weights."""
if lemma.endswith("ō"):
base = lemma[:-1]
# can't distinguish conj from 1sg alone reliably; default by ending vowel
if base.endswith("i"):
return ("3io", base[:-1], base[:-1] + "īv", base[:-1] + "īt")
return (3, base, base + "s", base + "t")
return None
# ── PUBLIC: participles ─────────────────────────────────────────────────────────
def participle(lemma, kind, case="nom", gender="m", number="singular"):
"""kind: 'prs' (present active, -ns/-ntis), 'pfv' (perfect passive, -tus),
'fut' (future active, -tūrus). Declined as an adjective via rule endings.
Returns (form, conf)."""
v = _VERBS.get(lemma)
if not v:
return lemma, "fallback"
conj, pstem, perfstem, supstem = v
if kind == "pfv":
if not supstem:
return lemma, "fallback"
base = supstem[:-1] if supstem.endswith("t") or supstem.endswith("s") else supstem
stem = supstem # supine stem already ends in t/s: amāt- -> amātus
return _decline_us_a_um(stem, case, gender, number), "rule"
if kind == "fut":
if not supstem:
return lemma, "fallback"
return _decline_us_a_um(supstem + "ūr", case, gender, number), "rule"
if kind == "prs":
# present active participle: stem + ns (nom), stem + nt- (oblique), 3rd-decl
pv = {1: "ā", 2: "ē", 3: "ē", "3io": "", 4: ""}[conj]
ntstem = pstem + pv + "nt"
return _decline_pres_ptcp(pstem + pv, case, gender, number), "rule"
return lemma, "fallback"
def _decline_us_a_um(stem, case, gender, number):
"""Decline a -us/-a/-um adjective/participle stem (2-1-2 declension)."""
C = _CASE_MAP.get(case, case.upper())
end = {
("NOM", "m", "singular"): "us", ("NOM", "f", "singular"): "a", ("NOM", "n", "singular"): "um",
("GEN", "m", "singular"): "ī", ("GEN", "f", "singular"): "ae", ("GEN", "n", "singular"): "ī",
("DAT", "m", "singular"): "ō", ("DAT", "f", "singular"): "ae", ("DAT", "n", "singular"): "ō",
("ACC", "m", "singular"): "um", ("ACC", "f", "singular"): "am", ("ACC", "n", "singular"): "um",
("ABL", "m", "singular"): "ō", ("ABL", "f", "singular"): "ā", ("ABL", "n", "singular"): "ō",
("VOC", "m", "singular"): "e", ("VOC", "f", "singular"): "a", ("VOC", "n", "singular"): "um",
("NOM", "m", "plural"): "ī", ("NOM", "f", "plural"): "ae", ("NOM", "n", "plural"): "a",
("GEN", "m", "plural"): "ōrum", ("GEN", "f", "plural"): "ārum", ("GEN", "n", "plural"): "ōrum",
("DAT", "m", "plural"): "īs", ("DAT", "f", "plural"): "īs", ("DAT", "n", "plural"): "īs",
("ACC", "m", "plural"): "ōs", ("ACC", "f", "plural"): "ās", ("ACC", "n", "plural"): "a",
("ABL", "m", "plural"): "īs", ("ABL", "f", "plural"): "īs", ("ABL", "n", "plural"): "īs",
("VOC", "m", "plural"): "ī", ("VOC", "f", "plural"): "ae", ("VOC", "n", "plural"): "a",
}.get((C, gender, number), "us")
return stem + end
def _decline_pres_ptcp(stem, case, gender, number):
"""Present active participle (amāns, amantis) — 3rd-declension, stem+ns/nt."""
C = _CASE_MAP.get(case, case.upper())
if C == "NOM" and number == "singular":
return stem + "ns"
if C == "VOC" and number == "singular":
return stem + "ns"
base = stem + "nt"
end = {
("GEN", "singular"): "is", ("DAT", "singular"): "ī",
("ACC", "singular"): "em" if gender != "n" else "",
("ABL", "singular"): "e",
("NOM", "plural"): "ēs" if gender != "n" else "ia",
("GEN", "plural"): "ium", ("DAT", "plural"): "ibus",
("ACC", "plural"): "ēs" if gender != "n" else "ia",
("ABL", "plural"): "ibus", ("VOC", "plural"): "ēs",
}.get((C, number), "is")
if C == "ACC" and number == "singular" and gender == "n":
return stem + "ns"
return base + end
def infinitive(lemma, tense="present", voice="active"):
lemma = lemma.strip()
if lemma == "sum":
return ("esse", "rule") if tense == "present" else ("fuisse", "rule")
v = _VERBS.get(lemma)
if not v:
return lemma, "fallback"
conj, pstem, perfstem, supstem = v
if tense == "present":
if voice == "active":
return _active_infinitive_stem(conj, pstem).rstrip() + \
("re" if conj != 3 and conj != "3io" else "re"), "rule"
# passive present infinitive
base = {1: pstem + "ā", 2: pstem + "ē", 4: pstem + "ī"}.get(conj)
if base:
return base + "", "rule"
return pstem + "ī", "rule" # 3rd: regī
if tense == "perfect" and voice == "active" and perfstem:
return perfstem + "isse", "rule"
return lemma, "fallback"
def lexicon_stats():
return {
"noun_adj_source": "UniMorph Latin (github.com/unimorph/lat, CC-BY-SA 3.0)",
"verb_source": "rule-based 4-conjugation engine over curated attested "
"principal parts (UniMorph verb list is a 947-lemma sample "
"MISSING all core verbs — amō/sum/videō absent)",
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
"curated_verb_lemmas": len(_VERBS) + len(_IRREG),
"gender_inference": "declension-based (nom+gen endings) + curated exceptions",
}
if __name__ == "__main__":
import json
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
print("\n-- noun declension puella (1st, fem) --")
for c in ("nom", "gen", "dat", "acc", "abl", "voc"):
print(f" {c}: sg={decline_noun('puella', c, 'singular')[0]:10} "
f"pl={decline_noun('puella', c, 'plural')[0]}")
print("\n-- rēx (3rd, m):", [decline_noun('rēx', c, 'singular')[0] for c in ('nom','gen','dat','acc','abl')])
print("-- gender: puella=", noun_gender("puella"), "rēx=", noun_gender("rēx"),
"bellum=", noun_gender("bellum"), "corpus=", noun_gender("corpus"),
"manus=", noun_gender("manus"), "diēs=", noun_gender("diēs"))
print("\n-- conjugate videō (2nd) present ind active --")
for p in ("first", "second", "third"):
for n in ("singular", "plural"):
print(f" {p[:3]}.{n[:2]}: {conjugate('videō','present','ind','active',p,n)[0]}")
print("-- amō forms:", conjugate("amō","present","ind","active","first","singular")[0],
conjugate("amō","imperfect","ind","active","third","plural")[0],
conjugate("amō","future","ind","active","first","singular")[0],
conjugate("amō","perfect","ind","active","third","singular")[0])
print("-- sum:", [conjugate("sum","present","ind","active",p,"singular")[0] for p in ("first","second","third")])
print("-- participle amō pfv acc.f.sg:", participle("amō","pfv","acc","f","singular")[0])
print("-- infinitive amō:", infinitive("amō")[0], "| regō pass:", infinitive("regō", voice="passive")[0])
-538
View File
@@ -1,538 +0,0 @@
"""morphology_pt_full.py — production-grade Brazilian-Portuguese morphological generator.
NOT a toy. Backed by two real, broad, Wiktionary-lineage lexicons:
VERBS UniMorph Portuguese (github.com/unimorph/por, CC-BY-SA 3.0)
4,001 verb lemmas × full paradigm (283,991 finite/non-finite forms +
20,005 participle forms). Every mood/tense pt actually inflects:
indicative present / preterite (PST;PFV) / imperfect (PST;IPFV) /
pluperfect-simple (PST;PRF) / future,
conditional (futuro do pretérito),
subjunctive present / imperfect / FUTURE (PT-specific live tense),
affirmative + negative imperative,
PERSONAL infinitive (V;{p};{n};NFIN a PT-specific finite-ish form),
past participle (4 gender/number forms) + gerúndio (V.PTCP;PRS).
NOUNS + ADJECTIVES kaikki.org Portuguese (Wiktionary extract, same lineage)
81,138 noun lemmas WITH inherent gender + real (often irregular) plural
so -ão-ões / -ãos / -ães / -õos is resolved PER LEMMA by Wiktionary,
never guessed (mãomãos, pãopães, coraçãocorações).
40,252 adjective lemmas with real feminine + masc/fem plural forms.
Fallbacks (degrade, never crash, on out-of-vocabulary input):
verbs : rule generator for regular -ar/-er/-ir paradigms
nouns : gender heuristic (endings) + rule pluralization (with -ão FLAGGED)
adjs : -o/-a gender rule + rule pluralization
Confidence flag on every form:
"lexicon" straight from UniMorph/kaikki (trust: high)
"rule" deterministic rule (trust: medium)
"fallback" could not inflect; returned lemma (trust: low -> FLAG)
Public API (used by realizer_pt.py):
conjugate(lemma, mood, tense, person, number) -> (form, conf)
personal_infinitive(lemma, person, number) -> (form, conf)
participle(lemma, gender="m", number="singular") -> (form, conf)
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"
inflect_noun(lemma, number, gender=None) -> (form, conf)
inflect_adj(lemma, gender, number) -> (form, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "por.unimorph")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_pt.jsonl")
_CACHE = os.path.join(_HERE, "data", "pt_morph_cache.pkl")
# ── mood/tense pair -> UniMorph feature triple (a in tag; b in tag; c in tag) ────
_VERB_KEYMAP = {
("ind", "present"): ("IND", "PRS", None),
("ind", "preterite"): ("IND", "PST", "PFV"),
("ind", "imperfect"): ("IND", "PST", "IPFV"),
("ind", "pluperfect"): ("IND", "PST", "PRF"), # simple mais-que-perfeito
("ind", "future"): ("IND", "FUT", None),
("ind", "conditional"): ("COND", None, None),
("sbjv", "present"): ("SBJV", "PRS", None),
("sbjv", "imperfect"): ("SBJV", "PST", "IPFV"),
("sbjv", "future"): ("SBJV", "FUT", None), # PT-specific
("imp", "affirmative"): ("IMP", "POS", None),
("imp", "negative"): ("IMP", "NEG", None),
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── build the compact lexicon from UniMorph (verbs) + kaikki (nouns/adjs) ────────
def _build_verbs():
verbs = {} # (lemma, "mood|tense|person|number") -> form
pinf = {} # (lemma, "person|number") -> personal-infinitive form
part = {} # lemma -> {("m","SG"): form, ...} past participle
ger = {} # lemma -> gerúndio
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f: # past participle: falado/falada/falados/faladas
g = "m" if "MASC" in f else ("f" if "FEM" in f else "m")
num = "SG" if "SG" in f else ("PL" if "PL" in f else "SG")
part.setdefault(lemma, {})[(g, num)] = form
elif "PRS" in f: # gerúndio: falando
ger.setdefault(lemma, form)
continue
if head != "V":
continue
# personal / impersonal infinitive
if "NFIN" in f:
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person and number:
pinf[(lemma, f"{person}|{number}")] = form
continue
# finite forms
mt = None
for (mood, tense), (a, b, c) in _VERB_KEYMAP.items():
if a not in f:
continue
if b is not None and b not in f:
continue
if c is not None and c not in f:
continue
# IND;PST needs exactly PFV|IPFV|PRF — reject if the required one absent
mt = (mood, tense)
break
if mt is None:
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
verbs.setdefault((lemma, f"{mt[0]}|{mt[1]}|{person}|{number}"), form)
return verbs, pinf, part, ger
def _kaikki_gender(arg):
if not arg:
return None
a = arg.lower()
if a.startswith("f"):
return "f"
if a.startswith("m"):
return "m"
return None
def _build_nouns_adjs():
nouns = {} # lemma -> {"g","SG","PL"}
adjs = {} # lemma -> {("m","SG"),("f","SG"),("m","PL"),("f","PL")}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
pos = d.get("pos")
word = d.get("word", "")
if not word or " " in word: # skip multiword entries
continue
forms = d.get("forms", []) or []
if pos == "noun":
ht = d.get("head_templates") or []
g = None
if ht:
g = _kaikki_gender((ht[0].get("args") or {}).get("1"))
if g is None:
tags = d.get("tags") or []
if "feminine" in tags:
g = "f"
elif "masculine" in tags:
g = "m"
pl = None
for x in forms:
t = x.get("tags") or []
if "plural" in t and "alternative" not in t and "obsolete" not in t:
pl = x.get("form")
break
# first entry wins; but a later entry with a plural fills a gap
if word not in nouns:
nouns[word] = {"g": g, "SG": word, "PL": pl}
else:
cur = nouns[word]
if cur.get("g") is None and g:
cur["g"] = g
if not cur.get("PL") and pl:
cur["PL"] = pl
elif pos == "adj":
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word)
for x in forms:
t = set(x.get("tags") or [])
fm = x.get("form")
if not fm or ("alternative" in t) or ("obsolete" in t):
continue
if "comparative" in t or "superlative" in t or \
"diminutive" in t or "augmentative" in t:
continue
if "feminine" in t and "plural" in t:
d0[("f", "PL")] = fm
elif "masculine" in t and "plural" in t:
d0[("m", "PL")] = fm
elif "feminine" in t:
d0[("f", "SG")] = fm
elif "plural" in t: # invariant-gender adj (feliz -> felizes)
d0[("m", "PL")] = d0.get(("m", "PL")) or fm
d0[("f", "PL")] = d0.get(("f", "PL")) or fm
return nouns, adjs
def _build_cache():
verbs, pinf, part, ger = _build_verbs()
nouns, adjs = _build_nouns_adjs()
data = {"verbs": verbs, "pinf": pinf, "part": part, "ger": ger,
"nouns": nouns, "adjs": adjs}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
newest_src = max(os.path.getmtime(_UNIMORPH),
os.path.getmtime(_KAIKKI) if os.path.exists(_KAIKKI) else 0)
if os.path.getmtime(_CACHE) >= newest_src:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PINF, _PART, _GER, _NOUNS, _ADJS = (
_LEX["verbs"], _LEX["pinf"], _LEX["part"], _LEX["ger"],
_LEX["nouns"], _LEX["adjs"])
# ── regular-ending rule fallback (deterministic, last resort) ────────────────────
def _vclass(lemma):
return lemma[-2:] if lemma[-2:] in ("ar", "er", "ir") else None
def _stem(lemma):
return lemma[:-2]
# endings indexed [1sg,2sg,3sg,1pl,2pl,3pl]
_REG = {
("ind", "present", "ar"): ["o", "as", "a", "amos", "ais", "am"],
("ind", "present", "er"): ["o", "es", "e", "emos", "eis", "em"],
("ind", "present", "ir"): ["o", "es", "e", "imos", "is", "em"],
("ind", "preterite", "ar"): ["ei", "aste", "ou", "amos", "astes", "aram"],
("ind", "preterite", "er"): ["i", "este", "eu", "emos", "estes", "eram"],
("ind", "preterite", "ir"): ["i", "iste", "iu", "imos", "istes", "iram"],
("ind", "imperfect", "ar"): ["ava", "avas", "ava", "ávamos", "áveis", "avam"],
("ind", "imperfect", "er"): ["ia", "ias", "ia", "íamos", "íeis", "iam"],
("ind", "imperfect", "ir"): ["ia", "ias", "ia", "íamos", "íeis", "iam"],
("sbjv", "present", "ar"): ["e", "es", "e", "emos", "eis", "em"],
("sbjv", "present", "er"): ["a", "as", "a", "amos", "ais", "am"],
("sbjv", "present", "ir"): ["a", "as", "a", "amos", "ais", "am"],
("sbjv", "imperfect", "ar"): ["asse", "asses", "asse", "ássemos", "ásseis", "assem"],
("sbjv", "imperfect", "er"): ["esse", "esses", "esse", "êssemos", "êsseis", "essem"],
("sbjv", "imperfect", "ir"): ["isse", "isses", "isse", "íssemos", "ísseis", "issem"],
("sbjv", "future", "ar"): ["ar", "ares", "ar", "armos", "ardes", "arem"],
("sbjv", "future", "er"): ["er", "eres", "er", "ermos", "erdes", "erem"],
("sbjv", "future", "ir"): ["ir", "ires", "ir", "irmos", "irdes", "irem"],
}
# future & conditional attach to the FULL infinitive
_FUT = ["ei", "ás", "á", "emos", "eis", "ão"]
_COND = ["ia", "ias", "ia", "íamos", "íeis", "iam"]
def _slot_idx(person, number):
base = {"first": 0, "second": 1, "third": 2}[person]
return base + (0 if number == "singular" else 3)
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
st, i = _stem(lemma), _slot_idx(person, number)
if mood == "ind" and tense == "future":
return lemma + _FUT[i]
if mood == "ind" and tense == "conditional":
return lemma + _COND[i]
if mood == "imp": # affirmative tú/vocês imperative ~ subjunctive present
table = _REG.get(("sbjv", "present", vc))
if table and tense == "negative":
return st + table[i]
# affirmative 2sg = 3sg present indicative; others = subjunctive
pres = _REG.get(("ind", "present", vc))
if person == "second" and number == "singular":
return st + pres[2]
return st + table[i] if table else None
table = _REG.get((mood, tense, vc))
if table:
return st + table[i]
return None
# verified corrections to UniMorph data errors (each audited individually, not
# guessed). The three 1PL-present entries are glued-allomorph errors surfaced by a
# full-lexicon scan for a non-final "mos" in V;1;PL;IND;PRS forms (the ONLY three).
_VERB_FIX = {
("estar", "ind", "imperfect", "third", "plural"): "estavam", # was "estávam"
("estar", "ind", "present", "first", "plural"): "estamos", # was "estamosestámos"
("haver", "ind", "present", "first", "plural"): "havemos", # was "havemoshemos"
("ir", "ind", "present", "first", "plural"): "vamos", # was "vamosimos"
}
# ── PUBLIC: verb conjugation ─────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
"""Return (surface, confidence). mood in ind|sbjv|imp; tense per _VERB_KEYMAP."""
lemma = lemma.strip().lower()
fix = _VERB_FIX.get((lemma, mood, tense, person, number))
if fix:
return fix, "lexicon"
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _VERBS.get((lemma, f"{mood}|{tense}|{p}|{n}"))
if form:
# pt-BR normalization: UniMorph `por` carries the EUROPEAN spelling of
# the -ar 1pl PRETERITE (-ámos). Brazilian PT drops the accent
# (falámos->falamos, chegámos->chegamos) — 3,334/4,001 verbs affected.
if (mood == "ind" and tense == "preterite" and person == "first"
and number == "plural" and form.endswith("ámos")):
form = form[:-4] + "amos"
return form, "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r:
return r, "rule"
return lemma, "fallback"
def personal_infinitive(lemma, person, number):
"""PT personal (inflected) infinitive: para falarmos, ao chegarem."""
lemma = lemma.strip().lower()
p, n = _PERSON.get(person), _NUMBER.get(number)
if p and n:
form = _PINF.get((lemma, f"{p}|{n}"))
if form:
return form, "lexicon"
# rule: infinitive + personal endings (-, -es, -, -mos, -des, -em)
end = {("first", "singular"): "", ("second", "singular"): "es",
("third", "singular"): "", ("first", "plural"): "mos",
("second", "plural"): "des", ("third", "plural"): "em"}.get((person, number), "")
return lemma + end, "rule"
# ── PUBLIC: participle + gerund ───────────────────────────────────────────────────
def participle(lemma, gender="m", number="singular"):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
d = _PART.get(lemma)
if d:
form = d.get((g, num)) or d.get(("m", "SG"))
if form:
return form, "lexicon"
if lemma.endswith("ar"):
base = lemma[:-2] + "ad"
elif lemma[-2:] in ("er", "ir"):
base = lemma[:-2] + "id"
else:
return lemma, "fallback"
suf = {"m|SG": "o", "f|SG": "a", "m|PL": "os", "f|PL": "as"}[f"{g}|{num}"]
return base + suf, "rule"
def gerund(lemma):
lemma = lemma.strip().lower()
if lemma in _GER:
return _GER[lemma], "lexicon"
if lemma.endswith("ar"):
return lemma[:-2] + "ando", "rule"
if lemma.endswith("er"):
return lemma[:-2] + "endo", "rule"
if lemma.endswith("ir"):
return lemma[:-2] + "indo", "rule"
return lemma, "fallback"
# ── PUBLIC: noun gender + number ─────────────────────────────────────────────────
_FEM_SUF = ("ção", "são", "ção", "dade", "tade", "agem", "igem", "ugem", "gem",
"ez", "eza", "ice", "ície", "tude", "ude", "âncbefore")
_FEM_SUF = ("ção", "são", "dade", "tade", "agem", "gem", "eza", "ez", "ice",
"tude", "ude", "ância", "ência", "ínia")
_MASC_SUF = ("ema", "oma", "ama", "grama", "eta", "ão") # Greek -ma etc. (mostly m)
def _gender_heuristic(noun):
for suf in _FEM_SUF:
if noun.endswith(suf):
return "f"
if noun.endswith(("ema", "oma", "ama")): # problema, idioma, programa
return "m"
if noun.endswith("a") or noun.endswith("ã"):
return "f"
if noun.endswith("o") or noun.endswith(("l", "r", "z", "m", "u", "i")):
return "m"
return "m"
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g"):
return d["g"]
return _gender_heuristic(lemma)
_INVARIANT_PL_SUF = ("s",) # paroxytones ending -s are invariant (o lápis / os lápis)
def _rule_plural(noun):
"""Deterministic PT pluralization. Returns (form, ok) where ok=False flags an
ambiguous -ão that should lower confidence (the lexicon normally resolves it)."""
if not noun:
return noun, True
if noun.endswith("ão"):
return noun[:-2] + "ões", False # majority rule, but AMBIGUOUS -> flag
if noun.endswith("m"):
return noun[:-1] + "ns", True # homem->homens, jardim->jardins
if noun.endswith("al"):
return noun[:-2] + "ais", True
if noun.endswith("el"):
return noun[:-2] + "éis", True
if noun.endswith("ol"):
return noun[:-2] + "óis", True
if noun.endswith("ul"):
return noun[:-2] + "uis", True
if noun.endswith("il"):
return noun[:-2] + "is", True # stressed (funil->funis); unstressed rarer
if noun.endswith(("r", "z")):
return noun + "es", True # flor->flores, luz->luzes
if noun.endswith("s"):
# paroxytone -s (lápis, ônibus) invariant; oxytone -s (país) -> -es
return noun, True
if noun.endswith(("a", "e", "i", "o", "u", "á", "é", "í", "ó", "ú", "ã")):
return noun + "s", True
return noun + "s", True
def inflect_noun(lemma, number, gender=None):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if number == "singular":
return (d["SG"] if d and d.get("SG") else lemma), ("lexicon" if d else "rule")
if d and d.get("PL"):
return d["PL"], "lexicon"
form, ok = _rule_plural(lemma)
return form, ("rule" if ok else "fallback")
# ── PUBLIC: adjective agreement ──────────────────────────────────────────────────
def inflect_adj(lemma, gender, number):
lemma = lemma.strip().lower()
g = "f" if gender == "f" else "m"
num = "SG" if number == "singular" else "PL"
d = _ADJS.get(lemma)
if d:
form = d.get((g, num))
if form:
return form, "lexicon"
# build a missing plural from this gender's singular
sg = d.get((g, "SG")) or d.get(("m", "SG")) or lemma
if num == "PL":
pl, ok = _rule_plural(sg)
return pl, ("rule" if ok else "fallback")
return sg, "lexicon"
# rule fallback: -o/-a gender, then pluralize
a = lemma
if g == "f":
if a.endswith("o"):
a = a[:-1] + "a"
elif a.endswith(("ês", "or")) and not a.endswith("ior"):
a = a + "a" # português->portuguesa, trabalhador->..a
if num == "PL":
a, ok = _rule_plural(a)
return a, ("rule" if ok else "fallback")
return a, "rule"
def lexicon_stats():
return {
"verb_source": "UniMorph Portuguese (github.com/unimorph/por)",
"noun_adj_source": "kaikki.org Portuguese (Wiktionary extract)",
"license": "CC-BY-SA (Wiktionary-derived)",
"verb_forms": len(_VERBS),
"verb_lemmas": len({k[0] for k in _VERBS}),
"personal_infinitive_forms": len(_PINF),
"participle_lemmas": len(_PART),
"gerund_lemmas": len(_GER),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
tests = [
("falar", "ind", "present", "first", "singular", "falo"),
("comer", "ind", "present", "third", "plural", "comem"),
("partir", "ind", "present", "first", "plural", "partimos"),
("ser", "ind", "present", "third", "singular", "é"),
("ir", "ind", "preterite", "first", "singular", "fui"),
("ter", "ind", "future", "first", "singular", "terei"),
("fazer", "sbjv", "present", "first", "singular", "faça"),
("dormir", "ind", "present", "first", "singular", "durmo"),
("dar", "ind", "preterite", "third", "singular", "deu"),
("poder", "ind", "conditional", "first", "singular", "poderia"),
("fazer", "sbjv", "future", "third", "singular", "fizer"),
("estar", "ind", "present", "third", "singular", "está"),
]
ok = 0
for lemma, mood, tense, per, num, exp in tests:
got, conf = conjugate(lemma, mood, tense, per, num)
flag = "OK " if got == exp else "XX "
ok += got == exp
print(f" {flag}{lemma:8} {mood}/{tense} {per[:3]}.{num[:2]} -> {got:14} ({conf}) exp={exp}")
print(f"verb tests {ok}/{len(tests)}")
print(" gender: casa=", noun_gender("casa"), "problema=", noun_gender("problema"),
"mão=", noun_gender("mão"), "coração=", noun_gender("coração"),
"flor=", noun_gender("flor"))
print(" plural: mão->", inflect_noun("mão", "plural"),
"| pão->", inflect_noun("pão", "plural"),
"| animal->", inflect_noun("animal", "plural"),
"| coração->", inflect_noun("coração", "plural"))
print(" adj: bonito/f/sg->", inflect_adj("bonito", "f", "singular"),
"| feliz/m/pl->", inflect_adj("feliz", "m", "plural"),
"| português/f/sg->", inflect_adj("português", "f", "singular"))
print(" part: fazer/m/sg->", participle("fazer"), "| ger falar->", gerund("falar"))
print(" pinf falar 1pl->", personal_infinitive("falar", "first", "plural"))
-609
View File
@@ -1,609 +0,0 @@
# -*- coding: utf-8 -*-
"""morphology_ro_full.py — production-grade Romanian morphological generator.
Romanian is the BIG typological delta of the Romance family. The verb engine and
the confidence/fallback contract TRANSFER from the Italian sibling; the NOMINAL
system is genuinely new: Romanian has a SUFFIXED definite article, a preserved
NOM/ACC vs GEN/DAT case distinction, a NEUTER gender (masc-agreeing in SG,
fem-agreeing in PL), and a VOCATIVE. Those are grounded in real per-lemma data,
not guessed.
Real, Wiktionary-lineage lexical sources:
VERBS UniMorph Romanian (github.com/unimorph/ron, CC-BY-SA 3.0)
~1216 verb lemmas × paradigm, CLEAN orthography:
indicativ prezent / imperfect (PST;IPFV) / perfectul simplu (PST;PFV) /
conjunctiv prezent (SBJV;PRS, stored WITHOUT the '' particle),
participiu (V.PTCP;PST, INVARIABLE in the perfect compus),
gerunziu (V.CVB;PRS), infinitiv (NFIN), imperativ.
ro_irreg_verbs (embedded) high-frequency verbs UniMorph MISSES
(avea, vrea, da) + the auxiliary clitic paradigms the compound tenses need
(perfect-compus am/ai/a/am/ați/au, viitor voi/vei/va/vom/veți/vor,
condițional /ai/ar/am/ați/ar). Real standard forms.
NOUNS kaikki.org Romanian (Wiktionary extract, CC-BY-SA 3.0)
the FULL declension per lemma, cleanly tagged:
(nom/acc | gen/dat | vocative) × (indefinite | definite) × (sg | pl).
This is what makes the suffixed article LEXICALLY grounded (omomul,
casăcasa, băiatbăiatul, casei gen/dat, omule vocative). Inherent gender
m / f / n (NEUTER available directly) from the head template.
ADJECTIVES UniMorph Romanian ADJ
full case × gender(MASC/FEM/NEUT) × number × definiteness paradigm.
Fallbacks (degrade, never crash, on OOV): rule verb conjugation for -a/-ea/-e/-i/-î
classes, rule pluralization, rule suffixed-article by gender+ending. Every form
carries a confidence flag: "lexicon" | "rule" | "fallback".
Public API (used by realizer_ro.py):
conjugate(lemma, mood, tense, person, number) -> (form, conf)
aux(kind, person, number) -> str # perfect / future / conditional clitics
participle(lemma) -> (form, conf) # INVARIABLE
gerund(lemma) -> (form, conf)
noun_gender(lemma) -> "m"|"f"|"n"
definite_suffix(noun, gender, number, case) -> (form, conf) # rule engine
inflect_noun(lemma, number, gender=None, case="nomacc", definite=False) -> (form, conf)
inflect_adj(lemma, gender, number, case="nomacc", definite=False) -> (form, conf)
lexicon_stats() -> dict
"""
import json
import os
import pickle
_HERE = os.path.dirname(os.path.abspath(__file__))
_UNIMORPH = os.path.join(_HERE, "data", "ron.unimorph")
_KAIKKI = os.path.join(_HERE, "data", "kaikki_ro.jsonl")
_CACHE = os.path.join(_HERE, "data", "ro_morph_cache.pkl")
# ── (mood, tense) -> UniMorph feature set ─────────────────────────────────────────
_VERB_KEYMAP = {
("ind", "present"): {"IND", "PRS"},
("ind", "imperfect"): {"IND", "PST", "IPFV"},
("ind", "perfect_s"): {"IND", "PST", "PFV"}, # perfectul simplu (regional/lit.)
("sbjv", "present"): {"SBJV", "PRS"},
("imp", "affirmative"): {"POS", "IMP"},
}
_PERSON = {"first": "1", "second": "2", "third": "3"}
_NUMBER = {"singular": "SG", "plural": "PL"}
def _feat_set(tag):
return set(tag.split(";"))
# ── high-frequency irregulars UniMorph misses + auxiliary clitic paradigms ────────
# Real standard Romanian forms (textbook paradigms).
_IRREG = {
"avea": {
"ind|present|1|SG": "am", "ind|present|2|SG": "ai", "ind|present|3|SG": "are",
"ind|present|1|PL": "avem", "ind|present|2|PL": "aveți", "ind|present|3|PL": "au",
"ind|imperfect|1|SG": "aveam", "ind|imperfect|2|SG": "aveai",
"ind|imperfect|3|SG": "avea", "ind|imperfect|1|PL": "aveam",
"ind|imperfect|2|PL": "aveați", "ind|imperfect|3|PL": "aveau",
"sbjv|present|3|SG": "aibă", "sbjv|present|3|PL": "aibă",
"sbjv|present|1|SG": "am", "sbjv|present|2|SG": "ai",
"sbjv|present|1|PL": "avem", "sbjv|present|2|PL": "aveți",
"part": "avut", "ger": "având",
},
"vrea": {
"ind|present|1|SG": "vreau", "ind|present|2|SG": "vrei", "ind|present|3|SG": "vrea",
"ind|present|1|PL": "vrem", "ind|present|2|PL": "vreți", "ind|present|3|PL": "vor",
"ind|imperfect|1|SG": "voiam", "ind|imperfect|3|SG": "voia",
"sbjv|present|3|SG": "vrea", "sbjv|present|3|PL": "vrea",
"part": "vrut", "ger": "vrând",
},
"da": {
"ind|present|1|SG": "dau", "ind|present|2|SG": "dai", "ind|present|3|SG": "",
"ind|present|1|PL": "dăm", "ind|present|2|PL": "dați", "ind|present|3|PL": "dau",
"ind|imperfect|1|SG": "dădeam", "ind|imperfect|3|SG": "dădea",
"sbjv|present|3|SG": "dea", "sbjv|present|3|PL": "dea",
"part": "dat", "ger": "dând",
},
"fi": { # a fi — present is in UniMorph but keep participle + subjunctive here
"part": "fost", "ger": "fiind",
"sbjv|present|1|SG": "fiu", "sbjv|present|2|SG": "fii", "sbjv|present|3|SG": "fie",
"sbjv|present|1|PL": "fim", "sbjv|present|2|PL": "fiți", "sbjv|present|3|PL": "fie",
"ind|imperfect|1|SG": "eram", "ind|imperfect|2|SG": "erai",
"ind|imperfect|3|SG": "era", "ind|imperfect|1|PL": "eram",
"ind|imperfect|2|PL": "erați", "ind|imperfect|3|PL": "erau",
},
}
# auxiliary clitic paradigms (person,number)->form
_AUX = {
"perfect": {("first", "singular"): "am", ("second", "singular"): "ai",
("third", "singular"): "a", ("first", "plural"): "am",
("second", "plural"): "ați", ("third", "plural"): "au"},
"future": {("first", "singular"): "voi", ("second", "singular"): "vei",
("third", "singular"): "va", ("first", "plural"): "vom",
("second", "plural"): "veți", ("third", "plural"): "vor"},
"conditional": {("first", "singular"): "", ("second", "singular"): "ai",
("third", "singular"): "ar", ("first", "plural"): "am",
("second", "plural"): "ați", ("third", "plural"): "ar"},
}
def aux(kind, person, number):
return _AUX[kind][(person, number)]
# ── build verb lexicon from UniMorph ──────────────────────────────────────────────
def _build_verbs():
verbs, part, ger = {}, {}, {}
with open(_UNIMORPH, encoding="utf-8") as fh:
for line in fh:
line = line.rstrip("\n")
if not line or "\t" not in line:
continue
parts = line.split("\t")
if len(parts) != 3:
continue
lemma, form, tag = parts
f = _feat_set(tag)
head = tag.split(";")[0]
if head == "V.PTCP":
if "PST" in f:
part.setdefault(lemma, form)
continue
if head == "V.CVB":
if "PRS" in f:
ger.setdefault(lemma, form)
continue
if head != "V":
continue
person = next((p for p in ("1", "2", "3") if p in f), None)
number = "SG" if "SG" in f else ("PL" if "PL" in f else None)
if person is None or number is None:
continue
# conjunctiv forms in UniMorph carry a leading 'să ' — strip it
surf = form
if surf.startswith(""):
surf = surf[3:]
for (mood, tense), req in _VERB_KEYMAP.items():
if not req <= f:
continue
if tense == "imperfect" and "PFV" in f:
continue
if tense == "perfect_s" and "IPFV" in f:
continue
# keep IND;PRS out of the PRF slot (mai-mult-ca-perfect etc. ignored)
if {"IND", "PRS"} <= req and "PRF" in f:
continue
verbs.setdefault((lemma, f"{mood}|{tense}|{person}|{number}"), surf)
break
return verbs, part, ger
# ── kaikki nouns: full declension paradigm per lemma ──────────────────────────────
_EXCL = {"alternative", "archaic", "obsolete", "regional", "dialectal", "rare",
"table-tags", "inflection-template", "error-unrecognized-form",
"diminutive", "augmentative", "informal"}
def _noun_key(tagset):
if tagset & _EXCL:
return None
if "vocative" in tagset:
case = "voc"
elif "genitive" in tagset or "dative" in tagset:
case = "gendat"
elif "nominative" in tagset or "accusative" in tagset:
case = "nomacc"
else:
return None
definite = "definite" in tagset and "indefinite" not in tagset
number = "PL" if "plural" in tagset else ("SG" if "singular" in tagset else None)
if number is None:
return None
return (case, definite, number)
def _build_nouns():
nouns = {} # lemma -> {"g":..., para:{(case,def,num):form}, "PL":plain_plural}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
if d.get("pos") != "noun":
continue
word = d.get("word", "")
if not word or " " in word:
continue
ht = d.get("head_templates") or []
g = None
if ht:
a = str((ht[0].get("args") or {}).get("1") or "").lower()
if a[:1] in ("m", "f", "n"):
g = a[:1]
entry = nouns.setdefault(word, {"g": g, "para": {}, "PL": None})
if entry["g"] is None and g:
entry["g"] = g
for x in (d.get("forms") or []):
fm = x.get("form")
tg = set(x.get("tags") or [])
if not fm or fm in ("-", "#", "") or " " in fm:
continue
if tg == {"plural"} and not entry["PL"]:
entry["PL"] = fm
k = _noun_key(tg)
if k and k not in entry["para"]:
entry["para"][k] = fm
return nouns
# ── adjectives from kaikki (UniMorph ron ADJ is sparse AND mis-tagged; kaikki is
# clean: the 4-form agreement pattern bun/bună/buni/bune). Neuter maps sg->masc,
# pl->fem, so 4 forms (m/f × SG/PL) fully cover it. ────────────────────────────
def _build_adjs():
adjs = {} # lemma -> {(gender,number): form} gender in {m,f}
with open(_KAIKKI, encoding="utf-8") as fh:
for line in fh:
try:
d = json.loads(line)
except Exception:
continue
if d.get("pos") != "adj":
continue
word = d.get("word", "")
if not word or " " in word:
continue
d0 = adjs.setdefault(word, {})
d0.setdefault(("m", "SG"), word) # masc sg = headword
for x in (d.get("forms") or []):
fm = x.get("form")
t = set(x.get("tags") or [])
if not fm or " " in fm or fm in ("-", "#") or (t & _EXCL):
continue
if "definite" in t or "genitive" in t or "dative" in t:
continue # keep indefinite nom/acc agr set
pl = "plural" in t
fem = "feminine" in t
masc = "masculine" in t
if fem and pl:
d0.setdefault(("f", "PL"), fm)
elif masc and pl:
d0.setdefault(("m", "PL"), fm)
elif fem and not pl:
d0.setdefault(("f", "SG"), fm)
elif pl and not fem and not masc: # bare plural -> both genders
d0.setdefault(("m", "PL"), fm)
d0.setdefault(("f", "PL"), fm)
return adjs
def _build_cache():
verbs, part, ger = _build_verbs()
nouns = _build_nouns()
adjs = _build_adjs()
data = {"verbs": verbs, "part": part, "ger": ger, "nouns": nouns, "adjs": adjs}
try:
with open(_CACHE, "wb") as fh:
pickle.dump(data, fh, protocol=pickle.HIGHEST_PROTOCOL)
except OSError:
pass
return data
def _load():
if os.path.exists(_CACHE):
srcs = [_UNIMORPH, _KAIKKI]
newest = max(os.path.getmtime(s) for s in srcs if os.path.exists(s))
if os.path.getmtime(_CACHE) >= newest:
try:
with open(_CACHE, "rb") as fh:
return pickle.load(fh)
except Exception:
pass
return _build_cache()
_LEX = _load()
_VERBS, _PART, _GER, _NOUNS, _ADJS = (
_LEX["verbs"], _LEX["part"], _LEX["ger"], _LEX["nouns"], _LEX["adjs"])
# ── rule verb conjugation fallback ────────────────────────────────────────────────
def _vclass(lemma):
if lemma.endswith("a"):
return "a"
if lemma.endswith("ea"):
return "ea"
if lemma.endswith("e"):
return "e"
if lemma.endswith("i"):
return "i"
if lemma.endswith("î"):
return "î"
return None
# regular present endings by class [1sg,2sg,3sg,1pl,2pl,3pl]
_REG_PRS = {
"a": ["", "i", "ă", "ăm", "ați", "ă"], # a lucra type (simplified)
"ea": ["", "i", "e", "em", "eți", "", ],
"e": ["", "i", "e", "em", "eți", ""],
"i": ["esc", "ești", "ește", "im", "iți", "esc"], # -i type (a vorbi)
"î": ["ăsc", "ăști", "ăște", "âm", "âți", "ăsc"],
}
_SLOT = {("first", "singular"): 0, ("second", "singular"): 1, ("third", "singular"): 2,
("first", "plural"): 3, ("second", "plural"): 4, ("third", "plural"): 5}
def _rule_conjugate(lemma, mood, tense, person, number):
vc = _vclass(lemma)
if vc is None:
return None
i = _SLOT[(person, number)]
body = lemma[:-len(vc)]
if mood == "ind" and tense == "present":
end = _REG_PRS[vc][i]
return body + end
if mood == "ind" and tense == "imperfect":
# -a/-i/-î -> stem + a/eai...; -e/-ea -> eam. Simplified regular imperfect.
stem = body
endings = {"a": ["am", "ai", "a", "am", "ați", "au"],
"i": ["eam", "eai", "ea", "eam", "eați", "eau"],
"î": ["am", "ai", "a", "am", "ați", "au"],
"e": ["eam", "eai", "ea", "eam", "eați", "eau"],
"ea": ["eam", "eai", "ea", "eam", "eați", "eau"]}[vc]
return stem + endings[i]
return None
# ── PUBLIC verb API ───────────────────────────────────────────────────────────────
def conjugate(lemma, mood, tense, person, number):
lemma = lemma.strip().lower()
key = f"{mood}|{tense}|{_PERSON.get(person,'?')}|{_NUMBER.get(number,'?')}"
ir = _IRREG.get(lemma)
if ir and key in ir:
return ir[key], "lexicon"
form = _VERBS.get((lemma, key))
if form:
return form, "lexicon"
r = _rule_conjugate(lemma, mood, tense, person, number)
if r is not None:
return r, "rule"
return lemma, "fallback"
def participle(lemma):
"""Past participle — INVARIABLE in the perfect compus (am mers, am văzut)."""
lemma = lemma.strip().lower()
ir = _IRREG.get(lemma)
if ir and "part" in ir:
return ir["part"], "lexicon"
if lemma in _PART:
return _PART[lemma], "lexicon"
vc = _vclass(lemma)
if vc == "a":
return lemma[:-1] + "at", "rule"
if vc in ("ea",):
return lemma[:-2] + "ut", "rule"
if vc == "i":
return lemma[:-1] + "it", "rule"
if vc == "î":
return lemma[:-1] + "ât", "rule"
if vc == "e":
return lemma[:-1] + "ut", "rule"
return lemma, "fallback"
def gerund(lemma):
lemma = lemma.strip().lower()
ir = _IRREG.get(lemma)
if ir and "ger" in ir:
return ir["ger"], "lexicon"
if lemma in _GER:
return _GER[lemma], "lexicon"
vc = _vclass(lemma)
if vc in ("a", "î"):
return lemma[:-1] + "ând", "rule"
if vc in ("ea", "e", "i"):
return lemma[:-len(vc)] + "ind", "rule"
return lemma, "fallback"
# ── noun gender ───────────────────────────────────────────────────────────────────
def noun_gender(lemma):
lemma = lemma.strip().lower()
d = _NOUNS.get(lemma)
if d and d.get("g") in ("m", "f", "n"):
return d["g"]
if lemma.endswith(("ă", "a", "e")):
return "f"
return "m"
# ── SUFFIXED DEFINITE ARTICLE — rule engine (fallback for OOV nouns) ───────────────
def definite_suffix(noun, gender, number, case="nomacc"):
"""Attach the enclitic definite article by gender + ending. Returns (form, conf).
This is the headline Romanian-specific engine extension."""
n = noun
g = gender
if number == "singular":
if g in ("m", "n"):
if case == "gendat":
# masc/neut gen-dat definite: -lui
if n.endswith("e"):
return n + "lui", "rule" # câine -> câinelui
if n.endswith("u"):
return n + "lui", "rule"
return n + "ului", "rule" # om -> omului
# nom/acc
if n.endswith("e"):
return n + "le", "rule" # câine -> câinele
if n.endswith("u"):
return n + "l", "rule" # codru -> codrul
if n.endswith("i"):
return n + "ul", "rule"
return n + "ul", "rule" # om -> omul
# feminine singular
if case == "gendat":
# fem gen/dat definite = plural-stem + i (casei, fetei) — needs plural;
# approximated as: -ă->-ei, -e->-ei, -a->-alei
if n.endswith("ă"):
return n[:-1] + "ei", "rule" # casă -> casei
if n.endswith("e"):
return n[:-1] + "ei", "rule" # carte -> cărții(approx cartei)
if n.endswith("a"):
return n[:-1] + "lei", "rule"
return n + "i", "rule"
# fem nom/acc
if n.endswith("ă"):
return n[:-1] + "a", "rule" # casă -> casa
if n.endswith("e"):
return n[:-1] + "ea", "rule" # carte -> cartea
if n.endswith("a"):
return n + "ua", "rule" # stea -> steaua
if n.endswith("i"):
return n + "a", "rule"
return n + "a", "rule"
# plural
if case == "gendat":
base = noun
return base + "lor", "rule" # -lor for all gen/dat pl
if g == "m":
return noun + "i", "rule" # oameni -> oamenii (+i)
return noun + "le", "rule" # case -> casele, trenuri->trenurile
# ── rule pluralization (fallback) ─────────────────────────────────────────────────
def _rule_plural(noun, gender):
if gender == "f":
if noun.endswith("ă"):
return noun[:-1] + "e"
if noun.endswith("e"):
return noun[:-1] + "i"
if noun.endswith("a"):
return noun[:-1] + "le"
return noun + "e"
if gender == "n":
return noun + "uri"
# masculine
if noun.endswith(("e",)):
return noun[:-1] + "i"
return noun + "i"
# ── PUBLIC noun inflection ────────────────────────────────────────────────────────
def inflect_noun(lemma, number, gender=None, case="nomacc", definite=False):
lemma = lemma.strip().lower()
g = gender or noun_gender(lemma)
d = _NOUNS.get(lemma)
numk = "SG" if number == "singular" else "PL"
if d:
if case == "voc":
form = d["para"].get(("voc", True, numk)) or d["para"].get(("voc", False, numk))
if form:
return form, "lexicon"
# try the exact paradigm cell from kaikki (lexically grounded)
form = d["para"].get((case, definite, numk))
if form:
return form, "lexicon"
# indefinite fallbacks from the paradigm
if not definite:
form = d["para"].get(("nomacc", False, numk))
if form:
return form, "lexicon"
if numk == "PL" and d.get("PL"):
return d["PL"], "lexicon"
if numk == "SG":
return lemma, "lexicon"
# rule path
base = lemma if number == "singular" else _rule_plural(lemma, g)
if definite:
return definite_suffix(base, g, number, case)
return base, ("rule" if d is None else "lexicon")
# ── PUBLIC adjective agreement ────────────────────────────────────────────────────
def _neuter_map(gender, number):
# neuter agrees masculine in SG, feminine in PL
if gender == "n":
return "m" if number == "singular" else "f"
return gender
def inflect_adj(lemma, gender, number, case="nomacc", definite=False):
lemma = lemma.strip().lower()
numk = "SG" if number == "singular" else "PL"
eg = _neuter_map(gender, number) # neuter -> masc(SG)/fem(PL)
d = _ADJS.get(lemma)
if d:
form = d.get((eg, numk))
if form:
return form, "lexicon"
# rule fallback: 4-form pattern bun/bună/buni/bune keyed by effective gender
a = lemma
if number == "singular":
if eg == "f":
if a.endswith("e"):
return a, "rule" # mare invariant sg
if a.endswith("u"):
return a[:-1] + "ă", "rule" # nou -> nouă
if a.endswith("ă"):
return a, "rule"
return a + "ă", "rule" # bun -> bună
return a, "rule" # masc/neut sg = lemma
# plural
if eg == "f":
if a.endswith("e"):
return a[:-1] + "i", "rule" # mare -> mari
if a.endswith("u"):
return a[:-1] + "e", "rule" # nou -> noue (approx; 'noi' irr)
if a.endswith("ă"):
return a[:-1] + "e", "rule"
return a + "e", "rule" # bun -> bune
# masc/neut(SG-only)->here masc pl -> -i
if a.endswith("e"):
return a[:-1] + "i", "rule" # mare -> mari
if a.endswith("u"):
return a[:-1] + "i", "rule"
return a + "i", "rule" # bun -> buni
def lexicon_stats():
return {
"verb_source": "UniMorph Romanian (github.com/unimorph/ron) + curated "
"irregulars (avea/vrea/da + aux clitic paradigms)",
"noun_source": "kaikki.org Romanian — full case/definite/vocative declension",
"adj_source": "UniMorph Romanian ADJ (case×gender×number×definiteness)",
"license": "CC-BY-SA 3.0 (Wiktionary/UniMorph lineage)",
"unimorph_verb_forms": len(_VERBS),
"unimorph_verb_lemmas": len({k[0] for k in _VERBS}),
"irregular_verb_lemmas": len(_IRREG),
"participle_lemmas": len(_PART),
"noun_lemmas": len(_NOUNS),
"adj_lemmas": len(_ADJS),
}
if __name__ == "__main__":
print(json.dumps(lexicon_stats(), indent=2, ensure_ascii=False))
print("\n── SUFFIXED DEFINITE ARTICLE (the headline delta) ──")
for n, g in [("om", "m"), ("băiat", "m"), ("casă", "f"), ("carte", "f"),
("tren", "n"), ("student", "m"), ("floare", "f")]:
sg = inflect_noun(n, "singular", g, "nomacc", True)
pl = inflect_noun(n, "plural", g, "nomacc", True)
gd = inflect_noun(n, "singular", g, "gendat", True)
vo = inflect_noun(n, "singular", g, "voc", False)
print(f" {n:8}({g}) def.sg={sg[0]:12} def.pl={pl[0]:14} "
f"gen/dat.sg={gd[0]:12} voc={vo[0]}")
print("\n── NEUTER split agreement (tren: masc SG / fem PL) ──")
print(" tren nou ->", inflect_noun("tren", "singular", "n")[0],
inflect_adj("nou", "n", "singular")[0])
print(" trenuri noi->", inflect_noun("tren", "plural", "n")[0],
inflect_adj("nou", "n", "plural")[0])
print("\n── verbs ──")
for l, m, t, p, n, in [("merge", "ind", "present", "third", "singular"),
("avea", "ind", "present", "first", "singular"),
("fi", "ind", "present", "third", "singular"),
("vorbi", "ind", "present", "third", "plural"),
("face", "sbjv", "present", "third", "singular"),
("lucra", "ind", "imperfect", "third", "singular")]:
print(f" {l:8}{m}/{t:10}{p[:3]}.{n[:2]} -> {conjugate(l,m,t,p,n)}")
print(" perfect-aux(3sg):", aux("perfect", "third", "singular"),
"| future(1sg):", aux("future", "first", "singular"),
"| cond(3sg):", aux("conditional", "third", "singular"))
print(" participle merge/vedea:", participle("merge"), participle("vedea"))
+1 -1
View File
@@ -81,7 +81,7 @@ jobs:
# Link to produce the engram binary
- name: Link engram binary
run: |
cc -std=c11 -O2 \
cc -std=c11 -O2 -DHAVE_CURL \
-I /usr/local/lib/el \
-o dist/engram \
dist/engram.c \
+1 -1
View File
@@ -88,7 +88,7 @@ jobs:
# Link to produce the engram binary
- name: Link engram binary
run: |
cc -std=c11 -O2 \
cc -std=c11 -O2 -DHAVE_CURL \
-I /usr/local/lib/el \
-o dist/engram \
dist/engram.c \
+1 -1
View File
@@ -62,7 +62,7 @@ jobs:
# Link to produce the engram binary
- name: Link engram binary
run: |
cc -std=c11 -O2 \
cc -std=c11 -O2 -DHAVE_CURL \
-I /usr/local/lib/el \
-o dist/engram \
dist/engram.c \
BIN
View File
Binary file not shown.
+35 -137
View File
@@ -10,8 +10,6 @@ el_val_t query_param(el_val_t path, el_val_t key);
el_val_t query_int(el_val_t path, el_val_t key, el_val_t default_val);
el_val_t extract_id(el_val_t path, el_val_t prefix);
el_val_t route_stats(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_act_stats(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_text_health(el_val_t method, el_val_t path, el_val_t body);
el_val_t persist_canonical(void);
el_val_t route_create_node(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_get_node(el_val_t method, el_val_t path, el_val_t body);
@@ -20,19 +18,16 @@ el_val_t route_scan_edges(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_search(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_activate(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_create_edge(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_create_edges_batch(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_neighbors(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_strengthen(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_forget(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_save(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_load(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_health(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_embed_backfill(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_sync(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_load_merge(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_emit_ise(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_capture_knowledge(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_similarity(el_val_t method, el_val_t path, el_val_t body);
el_val_t check_auth_ok(el_val_t method, el_val_t body);
el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body);
@@ -121,20 +116,11 @@ el_val_t route_stats(el_val_t method, el_val_t path, el_val_t body) {
return 0;
}
el_val_t route_act_stats(el_val_t method, el_val_t path, el_val_t body) {
return engram_act_stats_json();
return 0;
}
el_val_t route_text_health(el_val_t method, el_val_t path, el_val_t body) {
return engram_text_health_json();
return 0;
}
el_val_t persist_canonical(void) {
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_1 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_1 = (EL_STR("/tmp/engram")); } else { _if_result_1 = (dir_raw); } _if_result_1; });
return engram_save(el_str_concat(dir, EL_STR("/snapshot.json")));
engram_save(el_str_concat(dir, EL_STR("/snapshot.json")));
return 1;
return 0;
}
@@ -142,18 +128,9 @@ el_val_t route_create_node(el_val_t method, el_val_t path, el_val_t body) {
el_val_t content = json_get_string(body, EL_STR("content"));
el_val_t nt_raw = json_get_string(body, EL_STR("node_type"));
el_val_t node_type = ({ el_val_t _if_result_2 = 0; if (str_eq(nt_raw, EL_STR(""))) { _if_result_2 = (EL_STR("Memory")); } else { _if_result_2 = (nt_raw); } _if_result_2; });
el_val_t sal_present = json_get_raw(body, EL_STR("salience"));
el_val_t salience = ({ el_val_t _if_result_3 = 0; if (str_eq(sal_present, EL_STR(""))) { _if_result_3 = (el_from_float(0.5)); } else { _if_result_3 = (json_get_float(body, EL_STR("salience"))); } _if_result_3; });
el_val_t label_raw = json_get_string(body, EL_STR("label"));
el_val_t label = ({ el_val_t _if_result_4 = 0; if (str_eq(label_raw, EL_STR(""))) { _if_result_4 = (content); } else { _if_result_4 = (label_raw); } _if_result_4; });
el_val_t imp_present = json_get_raw(body, EL_STR("importance"));
el_val_t importance = ({ el_val_t _if_result_5 = 0; if (str_eq(imp_present, EL_STR(""))) { _if_result_5 = (el_from_float(0.5)); } else { _if_result_5 = (json_get_float(body, EL_STR("importance"))); } _if_result_5; });
el_val_t conf_present = json_get_raw(body, EL_STR("confidence"));
el_val_t confidence = ({ el_val_t _if_result_6 = 0; if (str_eq(conf_present, EL_STR(""))) { _if_result_6 = (el_from_float(1.0)); } else { _if_result_6 = (json_get_float(body, EL_STR("confidence"))); } _if_result_6; });
el_val_t tier_raw = json_get_string(body, EL_STR("tier"));
el_val_t tier = ({ el_val_t _if_result_7 = 0; if (str_eq(tier_raw, EL_STR(""))) { _if_result_7 = (EL_STR("Working")); } else { _if_result_7 = (tier_raw); } _if_result_7; });
el_val_t tags = json_get_string(body, EL_STR("tags"));
el_val_t id = engram_node_full(content, node_type, label, salience, importance, confidence, tier, tags);
el_val_t sal_raw = json_get_float(body, EL_STR("salience"));
el_val_t salience = ({ el_val_t _if_result_3 = 0; if ((sal_raw == el_from_float(0.0))) { _if_result_3 = (el_from_float(0.5)); } else { _if_result_3 = (sal_raw); } _if_result_3; });
el_val_t id = engram_node(content, node_type, salience);
el_val_t saved = persist_canonical();
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"id\":\""), id), EL_STR("\",\"content\":\"")), content), EL_STR("\",\"node_type\":\"")), node_type), EL_STR("\"}"));
return 0;
@@ -181,7 +158,7 @@ el_val_t route_scan_nodes(el_val_t method, el_val_t path, el_val_t body) {
el_val_t route_scan_edges(el_val_t method, el_val_t path, el_val_t body) {
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_8 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_8 = (EL_STR("/tmp/engram")); } else { _if_result_8 = (dir_raw); } _if_result_8; });
el_val_t dir = ({ el_val_t _if_result_4 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_4 = (EL_STR("/tmp/engram")); } else { _if_result_4 = (dir_raw); } _if_result_4; });
el_val_t snap_path = el_str_concat(dir, EL_STR("/.scan-export.json"));
engram_save(snap_path);
el_val_t snap = fs_read(snap_path);
@@ -197,22 +174,22 @@ el_val_t route_scan_edges(el_val_t method, el_val_t path, el_val_t body) {
}
el_val_t route_search(el_val_t method, el_val_t path, el_val_t body) {
el_val_t q = ({ el_val_t _if_result_9 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_9 = (query_param(path, EL_STR("q"))); } else { _if_result_9 = (json_get_string(body, EL_STR("query"))); } _if_result_9; });
el_val_t q = ({ el_val_t _if_result_5 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_5 = (query_param(path, EL_STR("q"))); } else { _if_result_5 = (json_get_string(body, EL_STR("query"))); } _if_result_5; });
el_val_t lim_url = query_int(path, EL_STR("limit"), 0);
el_val_t lim_body = json_get_int(body, EL_STR("limit"));
el_val_t lim_either = ({ el_val_t _if_result_10 = 0; if ((lim_url > 0)) { _if_result_10 = (lim_url); } else { _if_result_10 = (lim_body); } _if_result_10; });
el_val_t limit = ({ el_val_t _if_result_11 = 0; if ((lim_either > 0)) { _if_result_11 = (lim_either); } else { _if_result_11 = (20); } _if_result_11; });
el_val_t lim_either = ({ el_val_t _if_result_6 = 0; if ((lim_url > 0)) { _if_result_6 = (lim_url); } else { _if_result_6 = (lim_body); } _if_result_6; });
el_val_t limit = ({ el_val_t _if_result_7 = 0; if ((lim_either > 0)) { _if_result_7 = (lim_either); } else { _if_result_7 = (20); } _if_result_7; });
return engram_search_json(q, limit);
return 0;
}
el_val_t route_activate(el_val_t method, el_val_t path, el_val_t body) {
el_val_t q = ({ el_val_t _if_result_12 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_12 = (query_param(path, EL_STR("q"))); } else { _if_result_12 = (json_get_string(body, EL_STR("query"))); } _if_result_12; });
el_val_t q = ({ el_val_t _if_result_8 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_8 = (query_param(path, EL_STR("q"))); } else { _if_result_8 = (json_get_string(body, EL_STR("query"))); } _if_result_8; });
if (str_eq(q, EL_STR(""))) {
return err_json(EL_STR("missing query"));
}
el_val_t d_raw = ({ el_val_t _if_result_13 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_13 = (query_int(path, EL_STR("depth"), 3)); } else { _if_result_13 = (json_get_int(body, EL_STR("depth"))); } _if_result_13; });
el_val_t depth = ({ el_val_t _if_result_14 = 0; if ((d_raw > 0)) { _if_result_14 = (d_raw); } else { _if_result_14 = (3); } _if_result_14; });
el_val_t d_raw = ({ el_val_t _if_result_9 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_9 = (query_int(path, EL_STR("depth"), 3)); } else { _if_result_9 = (json_get_int(body, EL_STR("depth"))); } _if_result_9; });
el_val_t depth = ({ el_val_t _if_result_10 = 0; if ((d_raw > 0)) { _if_result_10 = (d_raw); } else { _if_result_10 = (3); } _if_result_10; });
return el_str_concat(el_str_concat(EL_STR("{\"results\":"), engram_activate_json(q, depth)), EL_STR("}"));
return 0;
}
@@ -221,50 +198,15 @@ el_val_t route_create_edge(el_val_t method, el_val_t path, el_val_t body) {
el_val_t from_id = json_get_string(body, EL_STR("from_id"));
el_val_t to_id = json_get_string(body, EL_STR("to_id"));
el_val_t rel_raw = json_get_string(body, EL_STR("relation"));
el_val_t relation = ({ el_val_t _if_result_15 = 0; if (str_eq(rel_raw, EL_STR(""))) { _if_result_15 = (EL_STR("associates")); } else { _if_result_15 = (rel_raw); } _if_result_15; });
el_val_t w_present = json_get_raw(body, EL_STR("weight"));
el_val_t weight = ({ el_val_t _if_result_16 = 0; if (str_eq(w_present, EL_STR(""))) { _if_result_16 = (el_from_float(0.5)); } else { _if_result_16 = (json_get_float(body, EL_STR("weight"))); } _if_result_16; });
el_val_t relation = ({ el_val_t _if_result_11 = 0; if (str_eq(rel_raw, EL_STR(""))) { _if_result_11 = (EL_STR("associates")); } else { _if_result_11 = (rel_raw); } _if_result_11; });
el_val_t w_raw = json_get_float(body, EL_STR("weight"));
el_val_t weight = ({ el_val_t _if_result_12 = 0; if ((w_raw == el_from_float(0.0))) { _if_result_12 = (el_from_float(0.5)); } else { _if_result_12 = (w_raw); } _if_result_12; });
engram_connect(from_id, to_id, weight, relation);
el_val_t saved = persist_canonical();
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"from_id\":\""), from_id), EL_STR("\",\"to_id\":\"")), to_id), EL_STR("\",\"relation\":\"")), relation), EL_STR("\"}"));
return 0;
}
el_val_t route_create_edges_batch(el_val_t method, el_val_t path, el_val_t body) {
el_val_t arr = json_get_raw(body, EL_STR("edges"));
if (str_eq(arr, EL_STR(""))) {
return err_json(EL_STR("missing edges array"));
}
el_val_t n = json_array_len(arr);
if (n == 0) {
return EL_STR("{\"ok\":true,\"accepted\":0,\"skipped\":0}");
}
el_val_t i = 0;
el_val_t accepted = 0;
el_val_t skipped = 0;
while (i < n) {
el_val_t item = json_array_get(arr, i);
el_val_t from_id = json_get_string(item, EL_STR("from_id"));
el_val_t to_id = json_get_string(item, EL_STR("to_id"));
if (str_eq(from_id, EL_STR("")) || str_eq(to_id, EL_STR(""))) {
skipped = (skipped + 1);
} else {
el_val_t rel_raw = json_get_string(item, EL_STR("relation"));
el_val_t relation = ({ el_val_t _if_result_17 = 0; if (str_eq(rel_raw, EL_STR(""))) { _if_result_17 = (EL_STR("associates")); } else { _if_result_17 = (rel_raw); } _if_result_17; });
el_val_t w_present = json_get_raw(item, EL_STR("weight"));
el_val_t weight = ({ el_val_t _if_result_18 = 0; if (str_eq(w_present, EL_STR(""))) { _if_result_18 = (el_from_float(0.5)); } else { _if_result_18 = (json_get_float(item, EL_STR("weight"))); } _if_result_18; });
engram_connect(from_id, to_id, weight, relation);
accepted = (accepted + 1);
}
i = (i + 1);
}
if (accepted > 0) {
el_val_t saved = persist_canonical();
}
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"accepted\":"), int_to_str(accepted)), EL_STR(",\"skipped\":")), int_to_str(skipped)), EL_STR("}"));
return 0;
}
el_val_t route_neighbors(el_val_t method, el_val_t path, el_val_t body) {
el_val_t id = extract_id(path, EL_STR("/api/neighbors/"));
if (str_eq(id, EL_STR(""))) {
@@ -300,51 +242,36 @@ el_val_t route_forget(el_val_t method, el_val_t path, el_val_t body) {
el_val_t route_save(el_val_t method, el_val_t path, el_val_t body) {
el_val_t p_raw = json_get_string(body, EL_STR("path"));
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_19 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_19 = (EL_STR("/tmp/engram")); } else { _if_result_19 = (dir_raw); } _if_result_19; });
el_val_t p = ({ el_val_t _if_result_20 = 0; if (str_eq(p_raw, EL_STR(""))) { _if_result_20 = (el_str_concat(dir, EL_STR("/snapshot.json"))); } else { _if_result_20 = (p_raw); } _if_result_20; });
el_val_t sv = engram_save(p);
el_val_t sv_ok = ({ el_val_t _if_result_21 = 0; if ((sv == 0)) { _if_result_21 = (EL_STR("false")); } else { _if_result_21 = (EL_STR("true")); } _if_result_21; });
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":"), sv_ok), EL_STR(",\"path\":\"")), p), EL_STR("\",\"node_count\":")), int_to_str(engram_node_count())), EL_STR(",\"edge_count\":")), int_to_str(engram_edge_count())), EL_STR("}"));
el_val_t dir = ({ el_val_t _if_result_13 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_13 = (EL_STR("/tmp/engram")); } else { _if_result_13 = (dir_raw); } _if_result_13; });
el_val_t p = ({ el_val_t _if_result_14 = 0; if (str_eq(p_raw, EL_STR(""))) { _if_result_14 = (el_str_concat(dir, EL_STR("/snapshot.json"))); } else { _if_result_14 = (p_raw); } _if_result_14; });
engram_save(p);
return el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"path\":\""), p), EL_STR("\"}"));
return 0;
}
el_val_t route_load(el_val_t method, el_val_t path, el_val_t body) {
el_val_t p_raw = json_get_string(body, EL_STR("path"));
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_22 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_22 = (EL_STR("/tmp/engram")); } else { _if_result_22 = (dir_raw); } _if_result_22; });
el_val_t p = ({ el_val_t _if_result_23 = 0; if (str_eq(p_raw, EL_STR(""))) { _if_result_23 = (el_str_concat(dir, EL_STR("/snapshot.json"))); } else { _if_result_23 = (p_raw); } _if_result_23; });
el_val_t ld = engram_load(p);
el_val_t ld_ok = ({ el_val_t _if_result_24 = 0; if ((ld == 0)) { _if_result_24 = (EL_STR("false")); } else { _if_result_24 = (EL_STR("true")); } _if_result_24; });
el_val_t nc_after = engram_node_count();
el_val_t hollow = ({ el_val_t _if_result_25 = 0; if ((nc_after == 0)) { _if_result_25 = (EL_STR("true")); } else { _if_result_25 = (EL_STR("false")); } _if_result_25; });
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":"), ld_ok), EL_STR(",\"path\":\"")), p), EL_STR("\",\"node_count\":")), int_to_str(nc_after)), EL_STR(",\"edge_count\":")), int_to_str(engram_edge_count())), EL_STR(",\"hollow\":")), hollow), EL_STR("}"));
el_val_t dir = ({ el_val_t _if_result_15 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_15 = (EL_STR("/tmp/engram")); } else { _if_result_15 = (dir_raw); } _if_result_15; });
el_val_t p = ({ el_val_t _if_result_16 = 0; if (str_eq(p_raw, EL_STR(""))) { _if_result_16 = (el_str_concat(dir, EL_STR("/snapshot.json"))); } else { _if_result_16 = (p_raw); } _if_result_16; });
engram_load(p);
return ok_json();
return 0;
}
el_val_t route_health(el_val_t method, el_val_t path, el_val_t body) {
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"status\":\"ok\",\"engine\":\"engram-runtime-native\",\"node_count\":"), int_to_str(engram_node_count())), EL_STR(",\"edge_count\":")), int_to_str(engram_edge_count())), EL_STR("}"));
return 0;
}
el_val_t route_embed_backfill(el_val_t method, el_val_t path, el_val_t body) {
el_val_t n = query_int(path, EL_STR("n"), 32);
el_val_t result = engram_embed_backfill(n);
el_val_t done = json_get_float(result, EL_STR("embedded"));
if (done > el_from_float(0.0)) {
el_val_t saved = persist_canonical();
}
return result;
return EL_STR("{\"status\":\"ok\",\"engine\":\"engram-runtime-native\"}");
return 0;
}
el_val_t route_sync(el_val_t method, el_val_t path, el_val_t body) {
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_26 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_26 = (EL_STR("/tmp/engram")); } else { _if_result_26 = (dir_raw); } _if_result_26; });
el_val_t dir = ({ el_val_t _if_result_17 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_17 = (EL_STR("/tmp/engram")); } else { _if_result_17 = (dir_raw); } _if_result_17; });
el_val_t snap_path = el_str_concat(dir, EL_STR("/.sync-export.json"));
engram_save(snap_path);
el_val_t snap = fs_read(snap_path);
if (str_eq(snap, EL_STR(""))) {
return err_json(EL_STR("sync export failed: snapshot unreadable"));
return EL_STR("{\"nodes\":[],\"edges\":[]}");
}
return snap;
return 0;
@@ -378,7 +305,7 @@ el_val_t route_emit_ise(el_val_t method, el_val_t path, el_val_t body) {
el_val_t conf = el_from_float(0.8);
el_val_t id = engram_node_full(content, EL_STR("InternalStateEvent"), EL_STR("state-event"), sal, imp, conf, EL_STR("Episodic"), EL_STR("[\"internal-state\",\"InternalStateEvent\"]"));
el_val_t ret_raw = env(EL_STR("ENGRAM_ISE_RETENTION_MS"));
el_val_t ret_ms = ({ el_val_t _if_result_27 = 0; if (str_eq(ret_raw, EL_STR(""))) { _if_result_27 = (172800000); } else { _if_result_27 = (str_to_int(ret_raw)); } _if_result_27; });
el_val_t ret_ms = ({ el_val_t _if_result_18 = 0; if (str_eq(ret_raw, EL_STR(""))) { _if_result_18 = (172800000); } else { _if_result_18 = (str_to_int(ret_raw)); } _if_result_18; });
el_val_t pruned = engram_prune_telemetry(ret_ms);
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"id\":\""), id), EL_STR("\",\"pruned\":")), int_to_str(pruned)), EL_STR("}"));
return 0;
@@ -390,21 +317,21 @@ el_val_t route_capture_knowledge(el_val_t method, el_val_t path, el_val_t body)
return err_json(EL_STR("missing content"));
}
el_val_t title = json_get_string(body, EL_STR("title"));
el_val_t label = ({ el_val_t _if_result_28 = 0; if (str_eq(title, EL_STR(""))) { _if_result_28 = (str_slice(content, 0, 60)); } else { _if_result_28 = (title); } _if_result_28; });
el_val_t label = ({ el_val_t _if_result_19 = 0; if (str_eq(title, EL_STR(""))) { _if_result_19 = (str_slice(content, 0, 60)); } else { _if_result_19 = (title); } _if_result_19; });
el_val_t category_raw = json_get_string(body, EL_STR("category"));
el_val_t category = ({ el_val_t _if_result_29 = 0; if (str_eq(category_raw, EL_STR(""))) { _if_result_29 = (EL_STR("other")); } else { _if_result_29 = (category_raw); } _if_result_29; });
el_val_t category = ({ el_val_t _if_result_20 = 0; if (str_eq(category_raw, EL_STR(""))) { _if_result_20 = (EL_STR("other")); } else { _if_result_20 = (category_raw); } _if_result_20; });
el_val_t ktier_raw = json_get_string(body, EL_STR("tier"));
el_val_t ktier = ({ el_val_t _if_result_30 = 0; if (str_eq(ktier_raw, EL_STR(""))) { _if_result_30 = (EL_STR("note")); } else { _if_result_30 = (ktier_raw); } _if_result_30; });
el_val_t ktier = ({ el_val_t _if_result_21 = 0; if (str_eq(ktier_raw, EL_STR(""))) { _if_result_21 = (EL_STR("note")); } else { _if_result_21 = (ktier_raw); } _if_result_21; });
el_val_t project = json_get_string(body, EL_STR("project"));
el_val_t tags_raw = json_get_raw(body, EL_STR("tags"));
el_val_t tags_base = ({ el_val_t _if_result_31 = 0; if (str_eq(tags_raw, EL_STR(""))) { _if_result_31 = (EL_STR("[]")); } else { _if_result_31 = (tags_raw); } _if_result_31; });
el_val_t tags_base = ({ el_val_t _if_result_22 = 0; if (str_eq(tags_raw, EL_STR(""))) { _if_result_22 = (EL_STR("[]")); } else { _if_result_22 = (tags_raw); } _if_result_22; });
el_val_t base_len = str_len(tags_base);
el_val_t head = str_slice(tags_base, 0, (base_len - 1));
el_val_t sep = ({ el_val_t _if_result_32 = 0; if (str_eq(head, EL_STR("["))) { _if_result_32 = (EL_STR("")); } else { _if_result_32 = (EL_STR(",")); } _if_result_32; });
el_val_t sep = ({ el_val_t _if_result_23 = 0; if (str_eq(head, EL_STR("["))) { _if_result_23 = (EL_STR("")); } else { _if_result_23 = (EL_STR(",")); } _if_result_23; });
el_val_t safe_cat = str_replace(category, EL_STR("\""), EL_STR("'"));
el_val_t safe_tier = str_replace(ktier, EL_STR("\""), EL_STR("'"));
el_val_t safe_proj = str_replace(project, EL_STR("\""), EL_STR("'"));
el_val_t proj_tag = ({ el_val_t _if_result_33 = 0; if (str_eq(safe_proj, EL_STR(""))) { _if_result_33 = (EL_STR("")); } else { _if_result_33 = (el_str_concat(el_str_concat(EL_STR(",\"project:"), safe_proj), EL_STR("\""))); } _if_result_33; });
el_val_t proj_tag = ({ el_val_t _if_result_24 = 0; if (str_eq(safe_proj, EL_STR(""))) { _if_result_24 = (EL_STR("")); } else { _if_result_24 = (el_str_concat(el_str_concat(EL_STR(",\"project:"), safe_proj), EL_STR("\""))); } _if_result_24; });
el_val_t tags = el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(head, sep), EL_STR("\"category:")), safe_cat), EL_STR("\",\"tier:")), safe_tier), EL_STR("\"")), proj_tag), EL_STR("]"));
el_val_t sal = el_from_float(0.5);
el_val_t imp = el_from_float(0.5);
@@ -415,20 +342,6 @@ el_val_t route_capture_knowledge(el_val_t method, el_val_t path, el_val_t body)
return 0;
}
el_val_t route_similarity(el_val_t method, el_val_t path, el_val_t body) {
el_val_t a = query_param(path, EL_STR("a"));
el_val_t b = query_param(path, EL_STR("b"));
if (str_eq(a, EL_STR(""))) {
return err_json(EL_STR("missing a"));
}
if (str_eq(b, EL_STR(""))) {
return err_json(EL_STR("missing b"));
}
el_val_t sim = engram_cosine_sim(a, b);
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"a\":\""), a), EL_STR("\",\"b\":\"")), b), EL_STR("\",\"cosine\":")), float_to_str(sim)), EL_STR("}"));
return 0;
}
el_val_t check_auth_ok(el_val_t method, el_val_t body) {
el_val_t key = env(EL_STR("ENGRAM_API_KEY"));
if (str_eq(key, EL_STR(""))) {
@@ -464,12 +377,6 @@ el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body) {
if (str_eq(method, EL_STR("GET")) && (str_eq(clean, EL_STR("/api/stats")) || str_eq(clean, EL_STR("/stats")))) {
return route_stats(method, path, body);
}
if (str_eq(method, EL_STR("GET")) && (str_eq(clean, EL_STR("/api/act-stats")) || str_eq(clean, EL_STR("/act-stats")))) {
return route_act_stats(method, path, body);
}
if (str_eq(method, EL_STR("GET")) && (str_eq(clean, EL_STR("/api/text-health")) || str_eq(clean, EL_STR("/text-health")))) {
return route_text_health(method, path, body);
}
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/nodes")) || str_eq(clean, EL_STR("/nodes")))) {
return route_create_node(method, path, body);
}
@@ -488,9 +395,6 @@ el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body) {
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/edges")) || str_eq(clean, EL_STR("/edges")))) {
return route_create_edge(method, path, body);
}
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/edges/batch")) || str_eq(clean, EL_STR("/edges/batch")))) {
return route_create_edges_batch(method, path, body);
}
if (str_eq(method, EL_STR("GET")) && str_starts_with(clean, EL_STR("/api/neighbors/"))) {
return route_neighbors(method, path, body);
}
@@ -521,12 +425,6 @@ el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body) {
if (str_eq(method, EL_STR("GET")) && str_eq(clean, EL_STR("/api/sync"))) {
return route_sync(method, path, body);
}
if (str_eq(clean, EL_STR("/api/embed-backfill"))) {
return route_embed_backfill(method, path, body);
}
if (str_eq(method, EL_STR("GET")) && str_starts_with(clean, EL_STR("/api/similarity"))) {
return route_similarity(method, path, body);
}
return el_str_concat(el_str_concat(EL_STR("{\"error\":\"not found\",\"path\":\""), clean), EL_STR("\"}"));
return 0;
}
@@ -534,10 +432,10 @@ el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body) {
int main(int _argc, char** _argv) {
el_runtime_init_args(_argc, _argv);
bind_raw = env(EL_STR("ENGRAM_BIND"));
bind_str = ({ el_val_t _if_result_34 = 0; if (str_eq(bind_raw, EL_STR(""))) { _if_result_34 = (EL_STR(":8742")); } else { _if_result_34 = (bind_raw); } _if_result_34; });
bind_str = ({ el_val_t _if_result_25 = 0; if (str_eq(bind_raw, EL_STR(""))) { _if_result_25 = (EL_STR(":8742")); } else { _if_result_25 = (bind_raw); } _if_result_25; });
port = parse_port(bind_str);
data_dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
data_dir = ({ el_val_t _if_result_35 = 0; if (str_eq(data_dir_raw, EL_STR(""))) { _if_result_35 = (EL_STR("/tmp/engram")); } else { _if_result_35 = (data_dir_raw); } _if_result_35; });
data_dir = ({ el_val_t _if_result_26 = 0; if (str_eq(data_dir_raw, EL_STR(""))) { _if_result_26 = (EL_STR("/tmp/engram")); } else { _if_result_26 = (data_dir_raw); } _if_result_26; });
snapshot_path = el_str_concat(data_dir, EL_STR("/snapshot.json"));
engram_load(snapshot_path);
boot_snap = fs_read(snapshot_path);
+13 -238
View File
@@ -76,38 +76,6 @@ fn route_stats(method: String, path: String, body: String) -> String {
engram_stats_json()
}
// route_act_stats GET /api/act-stats
// (2026-08-04 self-review) engram_act_stats_json() has existed since the
// 2026-07-27 review but was reachable ONLY through the soul daemon's heartbeat
// binding. Every activation-layer gauge WM evictions, breakthroughs, embedder
// breaker state, context drift, and now the Hebbian counters was therefore
// invisible unless the soul happened to be running and its ISEs were read back
// out of the store. Diagnosing the activation layer required a working soul,
// which is exactly backwards: the lower layer should be observable on its own.
// This review needed it to verify link formation and could not get at it. One
// line of plumbing, and the whole activation layer becomes directly diagnosable.
fn route_act_stats(method: String, path: String, body: String) -> String {
engram_act_stats_json()
}
// route_text_health GET /api/text-health
// (2026-08-08 self-review) The daily census half of the text-integrity gauge.
// Today's review found that the JSON parser had been replacing every \uXXXX
// escape with a literal '?' for at least two months: 3,119 of 4,081
// non-telemetry nodes (76%) were damaged, including the self traversal root
// and every values node, and NOTHING detected it because every gauge in the
// system measured whether the machinery was running, and none measured whether
// the text it carried was intact. No snapshot on disk predates the damage, so
// it cannot be undone; it can only be made impossible to repeat quietly.
//
// The parser is fixed. This route is the standing check: `damaged` should now
// hold flat at its historical floor and never climb. `write_damaged` (also on
// the heartbeat as txt_damaged) is the live regression signal non-zero means
// a write path is mangling text right now.
fn route_text_health(method: String, path: String, body: String) -> String {
engram_text_health_json()
}
// (2026-07-18 self-review) Scoping sweep: `let` inside an if-block creates an
// inner scope only it does NOT mutate the outer binding (documented with
// evidence in awareness.el, 2026-05-25). Every default/reassignment below used
@@ -133,54 +101,17 @@ fn route_text_health(method: String, path: String, body: String) -> String {
fn persist_canonical() -> Int {
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
// (2026-08-10 self-review) This returned a hardcoded 1, which made every
// caller's `let saved: Int = persist_canonical()` a dead variable six
// durable write paths each believed they had confirmation of a successful
// canonical persist and none of them had any. Propagate the real result.
return engram_save(dir + "/snapshot.json")
engram_save(dir + "/snapshot.json")
return 1
}
// INCOMPLETE-ROUTE FIX (2026-07-24 self-review): this route silently dropped
// label, importance, tier, and tags engram_node() defaults label to content
// and importance to 0.5, so every node created over HTTP lost its metadata.
// Observed live: the soul's boot-counter write-back landed with
// label="soul:boot_count:99" (content), importance 0.5, no tags. Honor the
// full field set via engram_node_full when any of them is supplied.
// PRESENCE-AWARE DEFAULTS (2026-08-01 self-review): the old pattern
// `if x == 0.0 { default }` made a legitimate 0.0 unrepresentable a caller
// setting salience/importance/weight to zero silently got 0.5. json_get_raw
// returns "" when the key is ABSENT and the raw token when present, so
// absence and zero are now distinguishable. Also: confidence was hardcoded
// to 1.0 regardless of input every HTTP-created node claimed full
// epistemic confidence. Now honored from the payload (default 1.0).
fn route_create_node(method: String, path: String, body: String) -> String {
let content: String = json_get_string(body, "content")
let nt_raw: String = json_get_string(body, "node_type")
let node_type: String = if str_eq(nt_raw, "") { "Memory" } else { nt_raw }
let sal_present: String = json_get_raw(body, "salience")
let salience: Float = if str_eq(sal_present, "") { 0.5 } else { json_get_float(body, "salience") }
let label_raw: String = json_get_string(body, "label")
let label: String = if str_eq(label_raw, "") { content } else { label_raw }
let imp_present: String = json_get_raw(body, "importance")
let importance: Float = if str_eq(imp_present, "") { 0.5 } else { json_get_float(body, "importance") }
let conf_present: String = json_get_raw(body, "confidence")
let confidence: Float = if str_eq(conf_present, "") { 1.0 } else { json_get_float(body, "confidence") }
let tier_raw: String = json_get_string(body, "tier")
let tier: String = if str_eq(tier_raw, "") { "Working" } else { tier_raw }
let tags: String = json_get_string(body, "tags")
// NO el_from_float WRAPPER (2026-08-01 self-review): salience/importance/
// confidence are already Float (el_val_t) values json_get_float and
// Float literals both encode. Wrapping them in el_from_float AGAIN
// reinterpreted the boxed bits as a raw double, producing garbage that
// failed engram_decode_score's range check and clamped every HTTP-created
// node to defaults (salience 0.9 in 0.5 stored; confidence 0.6 in → 1.0
// stored verified live). route_emit_ise always passed Floats bare and
// its 0.3/0.3/0.8 stored correctly; this call now does the same.
let id: String = engram_node_full(
content, node_type, label,
salience, importance, confidence,
tier, tags
)
let sal_raw: Float = json_get_float(body, "salience")
let salience: Float = if sal_raw == 0.0 { 0.5 } else { sal_raw }
let id: String = engram_node(content, node_type, salience)
let saved: Int = persist_canonical()
"{\"id\":\"" + id + "\",\"content\":\"" + content + "\",\"node_type\":\"" + node_type + "\"}"
}
@@ -246,64 +177,13 @@ fn route_create_edge(method: String, path: String, body: String) -> String {
let to_id: String = json_get_string(body, "to_id")
let rel_raw: String = json_get_string(body, "relation")
let relation: String = if str_eq(rel_raw, "") { "associates" } else { rel_raw }
// Presence-aware (2026-08-01): weight 0.0 is a legitimate edge weight
// (dormant association); only default when the key is absent.
let w_present: String = json_get_raw(body, "weight")
let weight: Float = if str_eq(w_present, "") { 0.5 } else { json_get_float(body, "weight") }
let w_raw: Float = json_get_float(body, "weight")
let weight: Float = if w_raw == 0.0 { 0.5 } else { w_raw }
engram_connect(from_id, to_id, weight, relation)
let saved: Int = persist_canonical()
"{\"ok\":true,\"from_id\":\"" + from_id + "\",\"to_id\":\"" + to_id + "\",\"relation\":\"" + relation + "\"}"
}
// route_create_edges_batch POST /api/edges/batch {"edges":[{from_id,to_id,relation,weight}, ...]}
//
// WHY THIS EXISTS (2026-08-07 self-review). persist_canonical() writes the
// FULL canonical snapshot 60MB at current graph size and route_create_edge
// calls it once per edge. That is correct for the interactive one-edge case and
// ruinous for any bulk write: the soul's Hebbian consolidation path delivers
// ~14 associations per 8-minute heartbeat, which through the single-edge route
// would be ~840MB of disk writes per beat, ~150GB/day, to persist 14 edges.
//
// The fix is not to weaken durability it is to make the unit of durability
// the BATCH. Connect every edge, then snapshot exactly once. Same guarantee
// (nothing acknowledged is lost to a restart), 1/N the writes. Empty or
// malformed entries are skipped rather than aborting the batch: a consolidation
// payload is best-effort by design, and one bad id should not cost the other 13.
//
// Returns the accepted count so the caller can tell delivery from silence.
fn route_create_edges_batch(method: String, path: String, body: String) -> String {
let arr: String = json_get_raw(body, "edges")
if str_eq(arr, "") { return err_json("missing edges array") }
let n: Int = json_array_len(arr)
if n == 0 { return "{\"ok\":true,\"accepted\":0,\"skipped\":0}" }
let i: Int = 0
let accepted: Int = 0
let skipped: Int = 0
while i < n {
let item: String = json_array_get(arr, i)
let from_id: String = json_get_string(item, "from_id")
let to_id: String = json_get_string(item, "to_id")
if str_eq(from_id, "") || str_eq(to_id, "") {
let skipped = skipped + 1
} else {
let rel_raw: String = json_get_string(item, "relation")
let relation: String = if str_eq(rel_raw, "") { "associates" } else { rel_raw }
let w_present: String = json_get_raw(item, "weight")
let weight: Float = if str_eq(w_present, "") { 0.5 } else { json_get_float(item, "weight") }
engram_connect(from_id, to_id, weight, relation)
let accepted = accepted + 1
}
let i = i + 1
}
// ONE snapshot for the whole batch the entire point of this route.
// Skip it when nothing was accepted: an all-malformed payload must not
// trigger a 60MB write.
if accepted > 0 {
let saved: Int = persist_canonical()
}
return "{\"ok\":true,\"accepted\":" + int_to_str(accepted) + ",\"skipped\":" + int_to_str(skipped) + "}"
}
fn route_neighbors(method: String, path: String, body: String) -> String {
let id: String = extract_id(path, "/api/neighbors/")
if str_eq(id, "") { return err_json("missing id") }
@@ -332,15 +212,8 @@ fn route_save(method: String, path: String, body: String) -> String {
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
let p: String = if str_eq(p_raw, "") { dir + "/snapshot.json" } else { p_raw }
// (2026-08-10 self-review) engram_save returns 0 on an empty path and the
// route discarded it, so the response was a literal "ok":true regardless
// of whether anything was written. Report the actual result AND the counts
// that were supposed to have been written the same move that made
// route_health honest on 2026-08-01. A caller can now tell "saved 13k
// nodes" from "saved nothing and said ok".
let sv: Int = engram_save(p)
let sv_ok: String = if sv == 0 { "false" } else { "true" }
"{\"ok\":" + sv_ok + ",\"path\":\"" + p + "\",\"node_count\":" + int_to_str(engram_node_count()) + ",\"edge_count\":" + int_to_str(engram_edge_count()) + "}"
engram_save(p)
"{\"ok\":true,\"path\":\"" + p + "\"}"
}
fn route_load(method: String, path: String, body: String) -> String {
@@ -348,59 +221,12 @@ fn route_load(method: String, path: String, body: String) -> String {
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
let p: String = if str_eq(p_raw, "") { dir + "/snapshot.json" } else { p_raw }
// (2026-08-10 self-review) This was a stub response over the single most
// destructive operation in the server. engram_load returns 0 on an empty
// path, an unopenable file, a zero-length file, or malloc failure and
// this route answered ok_json() in every one of those cases.
//
// Precise failure shape (el_runtime.c:9890): the fopen guard runs BEFORE
// the store reset, so a MISSING path is genuinely safe it returns 0 with
// the graph intact. The dangerous case is a readable-but-malformed file:
// the reset loop frees every node and edge FIRST, then parses, so a
// truncated or non-snapshot JSON leaves a hollow store and the caller
// was told "ok":true. With 37 GB of stale dated snapshots sitting in the
// data dir as tempting restore targets, "restore reported success and
// silently emptied the graph" is a live risk, not a hypothetical one.
//
// Fix: surface the return value AND the resulting counts. node_count=0
// after a load is the unambiguous hollow-store signal (same convention
// route_health adopted 2026-08-01). Callers can now verify a restore
// instead of trusting it.
let ld: Int = engram_load(p)
let ld_ok: String = if ld == 0 { "false" } else { "true" }
let nc_after: Int = engram_node_count()
let hollow: String = if nc_after == 0 { "true" } else { "false" }
"{\"ok\":" + ld_ok + ",\"path\":\"" + p + "\",\"node_count\":" + int_to_str(nc_after) + ",\"edge_count\":" + int_to_str(engram_edge_count()) + ",\"hollow\":" + hollow + "}"
engram_load(p)
ok_json()
}
// (2026-08-01 self-review) Health previously returned a hardcoded literal
// it reported "ok" even when the snapshot failed to load and the store was
// empty. Now reports live counts so a monitor can distinguish "up and
// loaded" from "up and hollow" (node_count=0 after boot = failed load).
fn route_health(method: String, path: String, body: String) -> String {
"{\"status\":\"ok\",\"engine\":\"engram-runtime-native\",\"node_count\":" + int_to_str(engram_node_count()) + ",\"edge_count\":" + int_to_str(engram_edge_count()) + "}"
}
// route_embed_backfill GET/POST /api/embed-backfill?n=48
//
// (2026-07-25 self-review) The lazy embedding backfill runs only inside
// engram_activate, and nothing in production calls /api/activate on this
// store the soul's curiosity loop activates its own in-process graph.
// After a restart from a snapshot without vectors, embedded_count stalled
// at 93/12175 and would never recover. This route lets the soul's
// heartbeat pump the backfill explicitly (48/min clears a 12k backlog in
// ~4h). Persists the canonical snapshot whenever new vectors were
// generated the 2026-07-25 regression happened precisely because 3747
// in-RAM embeddings were never snapshotted before a restart. Self-
// limiting: once coverage is full, embedded=0 and no save occurs.
fn route_embed_backfill(method: String, path: String, body: String) -> String {
let n: Int = query_int(path, "n", 32)
let result: String = engram_embed_backfill(n)
let done: Float = json_get_float(result, "embedded")
if done > 0.0 {
let saved: Int = persist_canonical()
}
return result
"{\"status\":\"ok\",\"engine\":\"engram-runtime-native\"}"
}
// route_sync return a snapshot of non-ISE/non-Working nodes for the soul daemon
@@ -423,16 +249,7 @@ fn route_sync(method: String, path: String, body: String) -> String {
let snap_path: String = dir + "/.sync-export.json"
engram_save(snap_path)
let snap: String = fs_read(snap_path)
// 2026-08-02 self-review: this used to return {"nodes":[],"edges":[]} when
// the export/read failed. The soul's sync_ok test (awareness.el) only
// checks for "" and "{}", so that placeholder PASSED as a healthy sync:
// soul.last_sync_ok_ts got stamped, sync_age_ms stayed green, the
// sync_empty warn ISE never fired, and engram_sync reported added:0
// forever. A totally broken sync was indistinguishable from a quiet
// healthy one the exact failure class this route was added to fix in
// the first place (see 2026-06-27 note above). Return a real error so the
// failure is loud on both sides.
if str_eq(snap, "") { return err_json("sync export failed: snapshot unreadable") }
if str_eq(snap, "") { return "{\"nodes\":[],\"edges\":[]}" }
return snap
}
@@ -554,25 +371,6 @@ fn route_capture_knowledge(method: String, path: String, body: String) -> String
"{\"ok\":true,\"id\":\"" + id + "\"}"
}
// route_similarity GET /api/similarity?a=<id>&b=<id>
//
// (2026-08-01 self-review) engram_cosine_sim was added 2026-07-24
// (bl-b2d1c944) with the stated purpose of exposing semantic distance to
// "EL code and the introspection API" but it had ZERO callers anywhere:
// no route, no soul-daemon use. The activation path uses embeddings
// internally (semantic seeding, Pass-2 additive term), but there was no way
// to probe pairwise node similarity from outside. This closes that: cosine
// in [-1,1], or -2 when either node is missing or not yet embedded (so
// "not comparable" is distinguishable from "genuinely orthogonal" 0.0).
fn route_similarity(method: String, path: String, body: String) -> String {
let a: String = query_param(path, "a")
let b: String = query_param(path, "b")
if str_eq(a, "") { return err_json("missing a") }
if str_eq(b, "") { return err_json("missing b") }
let sim: Float = engram_cosine_sim(a, b)
"{\"a\":\"" + a + "\",\"b\":\"" + b + "\",\"cosine\":" + float_to_str(sim) + "}"
}
// Auth
fn check_auth_ok(method: String, body: String) -> Bool {
@@ -619,12 +417,6 @@ fn handle_request(method: String, path: String, body: String) -> String {
if str_eq(method, "GET") && (str_eq(clean, "/api/stats") || str_eq(clean, "/stats")) {
return route_stats(method, path, body)
}
if str_eq(method, "GET") && (str_eq(clean, "/api/act-stats") || str_eq(clean, "/act-stats")) {
return route_act_stats(method, path, body)
}
if str_eq(method, "GET") && (str_eq(clean, "/api/text-health") || str_eq(clean, "/text-health")) {
return route_text_health(method, path, body)
}
// Nodes
if str_eq(method, "POST") && (str_eq(clean, "/api/nodes") || str_eq(clean, "/nodes")) {
@@ -647,13 +439,6 @@ fn handle_request(method: String, path: String, body: String) -> String {
if str_eq(method, "POST") && (str_eq(clean, "/api/edges") || str_eq(clean, "/edges")) {
return route_create_edge(method, path, body)
}
// Batch edge write one snapshot for the whole payload. Must be tested
// BEFORE nothing else claims it; the exact-match on "/api/edges" above
// does not catch "/api/edges/batch", so order is not load-bearing here,
// but keeping the two adjacent keeps them from drifting apart.
if str_eq(method, "POST") && (str_eq(clean, "/api/edges/batch") || str_eq(clean, "/edges/batch")) {
return route_create_edges_batch(method, path, body)
}
if str_eq(method, "GET") && str_starts_with(clean, "/api/neighbors/") {
return route_neighbors(method, path, body)
}
@@ -693,16 +478,6 @@ fn handle_request(method: String, path: String, body: String) -> String {
return route_sync(method, path, body)
}
// Embedding backfill pumped by the soul heartbeat (2026-07-25)
if str_eq(clean, "/api/embed-backfill") {
return route_embed_backfill(method, path, body)
}
// Semantic similarity probe (2026-08-01)
if str_eq(method, "GET") && str_starts_with(clean, "/api/similarity") {
return route_similarity(method, path, body)
}
"{\"error\":\"not found\",\"path\":\"" + clean + "\"}"
}
+10
View File
@@ -17,6 +17,16 @@
// 4. Append dep to order after all its transitive deps
// 5. Deduplicate: skip already-ordered vessels
// Cross-module forward declarations
// Defined in sibling epm modules; resolved at link time. The `extern fn` decls
// give elc the C prototypes so generated install.c compiles cleanly under strict
// compilers (gcc>=14 / clang) that reject implicit function declarations.
extern fn manifest_name(src: String) -> String // manifest.el
extern fn manifest_deps(src: String) -> String // manifest.el
extern fn registry_token() -> String // registry.el
extern fn registry_find(name: String, version: String) -> String // registry.el
extern fn registry_latest_version(name: String) -> String // registry.el
// Install paths
// packages_dir returns the root directory for installed vessels.
+9
View File
@@ -14,6 +14,15 @@
// EPM_REGISTRY_ORG org name that hosts vessel repos (default: neuron-technologies)
// EPM_TOKEN Gitea personal access token (required for publish)
// Cross-module forward declarations
// These symbols are defined in sibling epm modules or the El runtime and are
// resolved at link time. The `extern fn` decls give elc the C prototype so the
// generated registry.c compiles cleanly under strict compilers (gcc>=14 / clang)
// that reject implicit function declarations. Signature arity must match the
// definition; return/param types are informational (all lower to el_val_t).
extern fn config(key: String) -> String // El runtime builtin
extern fn read_installed() -> String // install.el
// Config helpers
// registry_api_url returns the Gitea API base URL with no trailing slash.
+9
View File
@@ -6,6 +6,15 @@
// Depends on: registry.el (registry_latest_version, registry_find),
// install.el (read_installed, install_vessel, installed_version)
// Cross-module forward declarations
// Defined in sibling epm modules; resolved at link time. The `extern fn` decls
// give elc the C prototypes so generated update.c compiles cleanly under strict
// compilers (gcc>=14 / clang) that reject implicit function declarations.
extern fn read_installed() -> String // install.el
extern fn installed_version(name: String) -> String // install.el
extern fn install_vessel(name: String, version: String) -> Bool // install.el
extern fn registry_latest_version(name: String) -> String // registry.el
// Semver helpers
// semver_part extracts the Nth dot-separated component from a semver string.
@@ -75,6 +75,7 @@ static inline void* el_win_dlsym(void* handle, const char* name) {
#include <direct.h> /* _mkdir */
#define mkdir(path, mode) _mkdir(path) /* POSIX mkdir(path,mode) → _mkdir(path) */
#define timegm _mkgmtime /* UTC tm → time_t */
#define fsync(fd) _commit(fd) /* no fsync() on Windows; _commit() (<io.h>) is the equiv */
/* setenv/unsetenv: not in the Windows CRT; map to _putenv_s / SetEnvironmentVariable. */
static inline int setenv(const char* name, const char* value, int overwrite) {
+380 -312
View File
@@ -82,8 +82,14 @@ static _Thread_local ElArena _tl_arena = {NULL, 0, 0};
static _Thread_local int _tl_arena_active = 0;
/* Binary-safe fs_read length — set by fs_read, consumed by http_send_response.
* Allows serving PNGs and other binary files without strlen truncation. */
static _Thread_local size_t _tl_fs_read_len = 0;
* Allows serving PNGs and other binary files without strlen truncation.
* PAIRED with the buffer pointer it describes: the length may only be applied
* to the exact buffer fs_read returned. Without the pairing, any handler that
* fs_read a file and then WRAPPED it into a larger response had that response
* truncated to the file's length (Content-Length lied AND the send stopped
* short) the safety-contact onboarding trap, 2026-07-17. */
static _Thread_local size_t _tl_fs_read_len = 0;
static _Thread_local const char* _tl_fs_read_buf = NULL;
static void el_arena_track(char* p) {
if (!_tl_arena_active || !p) return;
@@ -101,6 +107,8 @@ static void el_arena_track(char* p) {
void el_request_start(void) {
_tl_arena.count = 0;
_tl_arena_active = 1;
_tl_fs_read_len = 0; /* never let a previous request's file length */
_tl_fs_read_buf = NULL; /* leak into this response's byte accounting */
}
/* Called by http_worker after the El handler returns and the response is sent.
@@ -1484,11 +1492,14 @@ static void http_send_response(int fd, const char* body) {
}
const char* eff_body = is_envelope ? env_body : body;
/* Use the real byte count from fs_read if available (handles binary files
* with embedded null bytes PNG, WOFF2, etc.). Fall back to strlen for
* normal text/JSON responses where _tl_fs_read_len is 0. */
size_t blen = (_tl_fs_read_len > 0) ? _tl_fs_read_len : strlen(eff_body);
/* Use the real byte count from fs_read ONLY when this body IS the exact
* buffer fs_read returned (binary files with embedded null bytes PNG,
* WOFF2, etc.). Any other body wrapped, enveloped, or derived must be
* measured with strlen, or it is truncated/over-read to the file's size. */
size_t blen = (_tl_fs_read_len > 0 && eff_body == _tl_fs_read_buf)
? _tl_fs_read_len : strlen(eff_body);
_tl_fs_read_len = 0; /* consume — one-shot per response */
_tl_fs_read_buf = NULL;
int head_only = _tl_http_head_only;
JsonBuf hdrs; jb_init(&hdrs);
@@ -1545,17 +1556,6 @@ typedef struct {
#endif
} HttpWorkerArg;
/* Forward declarations for the loopback/API-key hardening helpers defined
* further down. Without these, http_worker's calls below were implicit
* declarations and the later `static` definitions conflicted with them this
* file did not compile at all. (2026-08-08 self-review: the hardening work
* they belong to had been sitting uncommitted in the working tree since
* 2026-07-15 in exactly this non-building state, which is presumably why it
* was never committed. Adding the two prototypes is the whole fix.) */
static int el_http_request_authorized(const char* method, const char* path,
const char* hdr_block);
static void el_http_send_401(int fd);
static void* http_worker(void* arg) {
HttpWorkerArg* a = (HttpWorkerArg*)arg;
#ifdef _WIN32
@@ -1564,13 +1564,8 @@ static void* http_worker(void* arg) {
int fd = a->fd;
#endif
free(a);
char *method = NULL, *path = NULL, *body = NULL, *hdr_block = NULL;
if (http_read_request(fd, &method, &path, &body, &hdr_block) == 0
&& !el_http_request_authorized(method, path, hdr_block)) {
/* Loopback hardening: EL_HTTP_AUTH_KEY is set and this request lacks the
* matching X-Neuron-Auth header refuse before it reaches any handler. */
el_http_send_401(fd);
} else if (method != NULL) {
char *method = NULL, *path = NULL, *body = NULL;
if (http_read_request(fd, &method, &path, &body, NULL) == 0) {
http_handler_fn h = http_lookup_active();
char* response = NULL;
/* HEAD: dispatch as GET so existing handlers respond with the same
@@ -1584,11 +1579,22 @@ static void* http_worker(void* arg) {
const char* rs = EL_CSTR(r);
/* Copy response out BEFORE arena teardown.
* For binary files, _tl_fs_read_len holds the real byte count
* use memcpy instead of strdup so null bytes are preserved. */
size_t rlen = _tl_fs_read_len > 0 ? _tl_fs_read_len : (rs ? strlen(rs) : 0);
* use memcpy instead of strdup so null bytes are preserved.
* The stored length applies ONLY when the response IS the exact
* fs_read buffer; a wrapped/derived response must use strlen or
* it gets truncated (or over-read) to the file's length. */
size_t rlen;
if (_tl_fs_read_len > 0 && rs && rs == _tl_fs_read_buf) {
rlen = _tl_fs_read_len; /* raw file bytes — binary-safe */
} else {
rlen = rs ? strlen(rs) : 0;
_tl_fs_read_len = 0; /* hint doesn't describe this body */
_tl_fs_read_buf = NULL;
}
response = malloc(rlen + 1);
if (response && rs) { memcpy(response, rs, rlen); response[rlen] = '\0'; }
else if (response) { response[0] = '\0'; }
if (_tl_fs_read_len > 0) _tl_fs_read_buf = response; /* hint follows the copy */
} else {
response = el_strdup_persist("el-runtime: no http handler registered");
}
@@ -1598,7 +1604,7 @@ static void* http_worker(void* arg) {
_tl_http_head_only = 0;
free(response);
}
free(method); free(path); free(body); free(hdr_block);
free(method); free(path); free(body);
el_closesocket(fd);
/* release a slot */
pthread_mutex_lock(&_http_conn_mu);
@@ -1608,108 +1614,6 @@ static void* http_worker(void* arg) {
return NULL;
}
/* ── loopback lock + local API-key auth (shipped desktop hardening) ────────
* Both controls are OFF by default (their env vars unset), so dev, self-host,
* and server builds behave exactly as before. The shipped macOS launcher
* neuron-daemons.sh sets them so a customer's soul is neither reachable from
* other machines on the LAN nor callable by other local users/processes
* without the per-install key held in the login Keychain:
*
* EL_HTTP_BIND_HOST=127.0.0.1 -> bind loopback only (el_http_apply_bind_addr)
* EL_HTTP_AUTH_KEY=<per-install> -> require "X-Neuron-Auth: <key>" per request
*/
/* Set the listen address on the dual-stack (AF_INET6, V6ONLY=0) socket. Default
* is in6addr_any (all interfaces) unchanged. When EL_HTTP_BIND_HOST names a
* loopback ("127.0.0.1", "localhost", "loopback", or "::1") we bind the IPv4-
* mapped IPv6 loopback ::ffff:127.0.0.1: on a V6ONLY=0 socket this accepts IPv4
* 127.0.0.1 clients (the desktop app connects there) while refusing every
* off-machine address. */
static void el_http_apply_bind_addr(struct sockaddr_in6* addr) {
const char* h = getenv("EL_HTTP_BIND_HOST");
int loopback = h && *h && (strcmp(h, "127.0.0.1") == 0
|| strcmp(h, "localhost") == 0
|| strcmp(h, "loopback") == 0
|| strcmp(h, "::1") == 0);
if (loopback) {
memset(&addr->sin6_addr, 0, sizeof(addr->sin6_addr));
addr->sin6_addr.s6_addr[10] = 0xff; /* ::ffff:127.0.0.1 */
addr->sin6_addr.s6_addr[11] = 0xff;
addr->sin6_addr.s6_addr[12] = 127;
addr->sin6_addr.s6_addr[15] = 1;
} else {
addr->sin6_addr = in6addr_any;
}
}
/* Human-readable description of the active bind host, for the listen log line. */
static const char* el_http_bind_desc(void) {
const char* h = getenv("EL_HTTP_BIND_HOST");
if (h && *h && (strcmp(h, "127.0.0.1") == 0 || strcmp(h, "localhost") == 0
|| strcmp(h, "loopback") == 0 || strcmp(h, "::1") == 0)) {
return "127.0.0.1 (loopback)";
}
return "[::] (dual-stack)";
}
/* Case-insensitive compare of the first n bytes of a and b. */
static int el_ci_eq_n(const char* a, const char* b, size_t n) {
for (size_t i = 0; i < n; i++) {
unsigned char ca = (unsigned char)a[i], cb = (unsigned char)b[i];
if (tolower(ca) != tolower(cb)) return 0;
}
return 1;
}
/* Return 1 iff the raw header block carries a header named `name` (case-
* insensitive) whose trimmed value equals `want` exactly. */
static int el_http_header_equals(const char* hdr_block, const char* name,
const char* want) {
if (!hdr_block || !name || !want) return 0;
size_t nlen = strlen(name), wlen = strlen(want);
const char* p = hdr_block;
while (*p) {
const char* line_end = strstr(p, "\r\n");
const char* end = line_end ? line_end : p + strlen(p);
const char* colon = memchr(p, ':', (size_t)(end - p));
if (colon && (size_t)(colon - p) == nlen && el_ci_eq_n(p, name, nlen)) {
const char* v = colon + 1;
while (v < end && (*v == ' ' || *v == '\t')) v++;
size_t vlen = (size_t)(end - v);
while (vlen > 0 && (v[vlen - 1] == ' ' || v[vlen - 1] == '\t')) vlen--;
if (vlen == wlen && memcmp(v, want, wlen) == 0) return 1;
}
if (!line_end) break;
p = line_end + 2;
}
return 0;
}
/* Authorize an inbound request. Enforcement is active only when EL_HTTP_AUTH_KEY
* is set; otherwise every request is allowed (dev default). GET/HEAD /health*
* are always allowed so launch-agent liveness probes work without the key. */
static int el_http_request_authorized(const char* method, const char* path,
const char* hdr_block) {
const char* key = getenv("EL_HTTP_AUTH_KEY");
if (!key || !*key) return 1;
if (method && (strcmp(method, "GET") == 0 || strcmp(method, "HEAD") == 0)
&& path && strncmp(path, "/health", 7) == 0) return 1;
return el_http_header_equals(hdr_block, "x-neuron-auth", key);
}
/* Minimal 401 for unauthorized requests — never reaches an EL handler. */
static void el_http_send_401(int fd) {
static const char* body = "{\"error\":\"unauthorized\",\"code\":\"auth_required\"}";
char resp[256];
int n = snprintf(resp, sizeof(resp),
"HTTP/1.1 401 Unauthorized\r\n"
"Content-Type: application/json\r\n"
"Content-Length: %zu\r\n"
"Connection: close\r\n\r\n%s",
strlen(body), body);
if (n > 0) http_send_all(fd, resp, (size_t)n);
}
el_val_t http_serve(el_val_t port, el_val_t handler) {
/* If `handler` looks like a string name, register it as the active handler. */
const char* hname = EL_CSTR(handler);
@@ -1728,13 +1632,13 @@ el_val_t http_serve(el_val_t port, el_val_t handler) {
struct sockaddr_in6 addr;
memset(&addr, 0, sizeof(addr));
addr.sin6_family = AF_INET6;
el_http_apply_bind_addr(&addr);
addr.sin6_addr = in6addr_any;
addr.sin6_port = htons((uint16_t)p);
if (bind(sock, (struct sockaddr*)&addr, sizeof(addr)) < 0) {
perror("bind"); el_closesocket(sock); return 0;
}
if (listen(sock, 64) < 0) { perror("listen"); el_closesocket(sock); return 0; }
fprintf(stderr, "[http] listening on %s port %d\n", el_http_bind_desc(), p);
fprintf(stderr, "[http] listening on [::]:%d (dual-stack)\n", p);
while (1) {
struct sockaddr_in6 cli;
socklen_t clen = sizeof(cli);
@@ -1940,10 +1844,20 @@ static void* http_worker_v2(void* arg) {
el_val_t hmap = http_build_headers_map(hdr_block ? hdr_block : "");
el_val_t r = h(EL_STR(dispatch_method), EL_STR(path), hmap, EL_STR(body));
const char* rs = EL_CSTR(r);
size_t rlen = _tl_fs_read_len > 0 ? _tl_fs_read_len : (rs ? strlen(rs) : 0);
/* Same pairing rule as the v1 worker: the fs_read length is only
* trustworthy for the exact buffer fs_read returned. */
size_t rlen;
if (_tl_fs_read_len > 0 && rs && rs == _tl_fs_read_buf) {
rlen = _tl_fs_read_len; /* raw file bytes — binary-safe */
} else {
rlen = rs ? strlen(rs) : 0;
_tl_fs_read_len = 0; /* hint doesn't describe this body */
_tl_fs_read_buf = NULL;
}
response = malloc(rlen + 1);
if (response && rs) { memcpy(response, rs, rlen); response[rlen] = '\0'; }
else if (response) { response[0] = '\0'; }
if (_tl_fs_read_len > 0) _tl_fs_read_buf = response; /* hint follows the copy */
el_release(hmap);
} else {
response = el_strdup_persist(
@@ -1984,13 +1898,13 @@ el_val_t http_serve_v2(el_val_t port, el_val_t handler) {
struct sockaddr_in6 addr;
memset(&addr, 0, sizeof(addr));
addr.sin6_family = AF_INET6;
el_http_apply_bind_addr(&addr);
addr.sin6_addr = in6addr_any;
addr.sin6_port = htons((uint16_t)p);
if (bind(sock, (struct sockaddr*)&addr, sizeof(addr)) < 0) {
perror("bind"); el_closesocket(sock); return 0;
}
if (listen(sock, 64) < 0) { perror("listen"); el_closesocket(sock); return 0; }
fprintf(stderr, "[http v2] listening on %s port %d\n", el_http_bind_desc(), p);
fprintf(stderr, "[http v2] listening on [::]:%d (dual-stack)\n", p);
while (1) {
struct sockaddr_in6 cli;
socklen_t clen = sizeof(cli);
@@ -2081,18 +1995,19 @@ void http_serve_async(el_val_t port, el_val_t handler) {
int sock = socket(AF_INET6, SOCK_STREAM, 0);
if (sock < 0) { perror("socket"); return; }
int yes = 1; int no = 0;
setsockopt(sock, SOL_SOCKET, SO_REUSEADDR, &yes, sizeof(yes));
setsockopt(sock, IPPROTO_IPV6, IPV6_V6ONLY, &no, sizeof(no));
/* Win32/mingw setsockopt takes optval as (const char*); the cast is portable on POSIX too. */
setsockopt(sock, SOL_SOCKET, SO_REUSEADDR, (const char*)&yes, sizeof(yes));
setsockopt(sock, IPPROTO_IPV6, IPV6_V6ONLY, (const char*)&no, sizeof(no));
struct sockaddr_in6 addr;
memset(&addr, 0, sizeof(addr));
addr.sin6_family = AF_INET6;
el_http_apply_bind_addr(&addr);
addr.sin6_addr = in6addr_any;
addr.sin6_port = htons((uint16_t)p);
if (bind(sock, (struct sockaddr*)&addr, sizeof(addr)) < 0) {
perror("bind"); close(sock); return;
}
if (listen(sock, 64) < 0) { perror("listen"); close(sock); return; }
fprintf(stderr, "[http] async listening on %s port %d\n", el_http_bind_desc(), p);
fprintf(stderr, "[http] async listening on [::]:%d (dual-stack)\n", p);
HttpServeAsyncArg* a = malloc(sizeof(HttpServeAsyncArg));
if (!a) { close(sock); return; }
a->sock = sock;
@@ -2141,6 +2056,7 @@ el_val_t http_response(el_val_t status, el_val_t headers_json, el_val_t body) {
el_val_t fs_read(el_val_t pathv) {
const char* path = EL_CSTR(pathv);
_tl_fs_read_len = 0;
_tl_fs_read_buf = NULL;
if (!path) return el_wrap_str(el_strdup(""));
FILE* f = fopen(path, "rb");
if (!f) return el_wrap_str(el_strdup(""));
@@ -2152,6 +2068,7 @@ el_val_t fs_read(el_val_t pathv) {
size_t got = fread(buf, 1, (size_t)sz, f);
buf[got] = '\0';
_tl_fs_read_len = got; /* store real byte count for binary-safe send */
_tl_fs_read_buf = buf; /* ...valid ONLY for this exact buffer */
fclose(f);
return el_wrap_str(buf);
}
@@ -3257,72 +3174,10 @@ static char* jp_parse_string_raw(JsonParser* jp) {
case 'r': c = '\r'; break;
case 't': c = '\t'; break;
case 'u': {
/* Decode \uXXXX (with surrogate pairs) to UTF-8.
* Ported from lang/releases/v1.0.0-20260501 (2026-08-08
* self-review). This copy carried the identical defect:
* the escape was skipped and a literal '?' emitted, which
* silently destroyed every non-ASCII character in any JSON
* string entering the runtime. Two copies of one parser
* bug is exactly how this class of fault survives, so the
* fix lands in both. See the release copy for the full
* measurement and rationale. */
unsigned cp = 0;
int ok = 1;
for (int i = 0; i < 4; i++) {
if (jp->p >= jp->end) { ok = 0; break; }
char h = *jp->p++;
unsigned d;
if (h >= '0' && h <= '9') d = (unsigned)(h - '0');
else if (h >= 'a' && h <= 'f') d = (unsigned)(h - 'a' + 10);
else if (h >= 'A' && h <= 'F') d = (unsigned)(h - 'A' + 10);
else { ok = 0; break; }
cp = (cp << 4) | d;
}
if (!ok) { c = '?'; break; }
if (cp >= 0xD800 && cp <= 0xDBFF &&
(size_t)(jp->end - jp->p) >= 6 &&
jp->p[0] == '\\' && jp->p[1] == 'u') {
const char* save = jp->p;
unsigned lo = 0; int ok2 = 1;
jp->p += 2;
for (int i = 0; i < 4; i++) {
char h = *jp->p++;
unsigned d;
if (h >= '0' && h <= '9') d = (unsigned)(h - '0');
else if (h >= 'a' && h <= 'f') d = (unsigned)(h - 'a' + 10);
else if (h >= 'A' && h <= 'F') d = (unsigned)(h - 'A' + 10);
else { ok2 = 0; break; }
lo = (lo << 4) | d;
}
if (ok2 && lo >= 0xDC00 && lo <= 0xDFFF)
cp = 0x10000u + ((cp - 0xD800u) << 10) + (lo - 0xDC00u);
else jp->p = save;
}
if (cp >= 0xD800 && cp <= 0xDFFF) cp = 0xFFFD;
char ub[4]; int un;
if (cp < 0x80) {
ub[0] = (char)cp; un = 1;
} else if (cp < 0x800) {
ub[0] = (char)(0xC0 | (cp >> 6));
ub[1] = (char)(0x80 | (cp & 0x3F)); un = 2;
} else if (cp < 0x10000) {
ub[0] = (char)(0xE0 | (cp >> 12));
ub[1] = (char)(0x80 | ((cp >> 6) & 0x3F));
ub[2] = (char)(0x80 | (cp & 0x3F)); un = 3;
} else {
ub[0] = (char)(0xF0 | (cp >> 18));
ub[1] = (char)(0x80 | ((cp >> 12) & 0x3F));
ub[2] = (char)(0x80 | ((cp >> 6) & 0x3F));
ub[3] = (char)(0x80 | (cp & 0x3F)); un = 4;
}
while (len + (size_t)un >= cap) {
cap *= 2;
out = realloc(out, cap);
if (!out) { fputs("el_runtime: out of memory\n", stderr); exit(1); }
}
for (int i = 0; i < un; i++) out[len++] = ub[i];
continue; /* bytes already appended */
/* Skip 4 hex digits; emit '?' as a placeholder */
for (int i = 0; i < 4 && jp->p < jp->end; i++) jp->p++;
c = '?';
break;
}
default: c = esc; break;
}
@@ -3756,8 +3611,10 @@ el_val_t json_get_raw(el_val_t json_str, el_val_t key) {
const char* k = EL_CSTR(key);
const char* p = json_find_key(json, k);
/* Clear fs_read binary-length hint — result is a fresh null-terminated
* string, not the raw file bytes, so Content-Length must use strlen. */
* string, not the raw file bytes, so Content-Length must use strlen.
* (Kept although the pointer pairing now makes this redundant.) */
_tl_fs_read_len = 0;
_tl_fs_read_buf = NULL;
if (!p) return el_wrap_str(el_strdup(""));
const char* end = json_skip_value(p);
size_t n = (size_t)(end - p);
@@ -6228,13 +6085,6 @@ void el_cgi_init(el_val_t name, el_val_t dharma_id, el_val_t principal,
#define ENGRAM_SUPPRESSION_BREAKTHROUGH 5
#define ENGRAM_BREAKTHROUGH_WEIGHT 0.25
#define ENGRAM_INHIBITION_FACTOR 0.1
/* ENGRAM_WM_CAP: hard global ceiling on nodes holding working_memory_weight
* > 0 at any time. Cowan (2001) puts human WM capacity at ~4 chunks; 24 gives
* the daemon generous headroom while preventing the unbounded growth observed
* in production (wm_active 288-778 per heartbeat "working memory" that is
* really the whole recently-touched graph). Ported from release runtime
* v1.0.0-20260501 Pass 5 on 2026-07-15 self-review. */
#define ENGRAM_WM_CAP 24
/* ── Layered consciousness architecture ──────────────────────────────────────
*
@@ -7082,6 +6932,243 @@ static int engram_rank_cmp(const void* a, const void* b) {
return 0;
}
/* ══════════════════════════════════════════════════════════════════════════
* SEMANTIC SEARCH LAYER nomic-embed-text via Ollama /api/embeddings
*
* Augments the lexical (istr_contains) matcher with dense-vector retrieval.
* Node content and the query are embedded through a local Ollama server;
* nodes are ranked by cosine similarity and UNIONED with lexical hits. This
* lets a paraphrase query surface a node whose words never appear in it.
*
* DEGRADABLE BY DESIGN. The whole layer is gated on HAVE_CURL plus a one-shot
* runtime probe of the embedding endpoint. If curl is not compiled in, or
* Ollama is unreachable, or ENGRAM_SEMANTIC=0, every entry point returns
* "no semantic signal" and callers fall back to pure lexical behaviour
* byte-for-byte the pre-existing search.
*
* CACHE. Node embeddings are computed lazily on first use and cached in
* process memory keyed by node id, with an FNV-1a content hash for
* invalidation (edited content re-embeds). The query is embedded once per
* search call. This is what "avoid re-embedding the whole graph every query"
* buys us: a warm cache serves cosine from RAM. (A cold process still pays
* O(N) embed calls the first time each node is scanned persisting the cache
* to a snapshot sidecar is the documented next step, not done here.)
*
* nomic task prefixes ("search_query:" / "search_document:") are applied
* because nomic-embed-text is trained with them; they materially improve
* retrieval separation (empirically: paraphrase 0.72 vs distractors <0.48).
*
* ENV:
* ENGRAM_SEMANTIC "0" disables; unset/other = auto-probe
* ENGRAM_EMBED_URL default http://localhost:11434/api/embeddings
* ENGRAM_EMBED_MODEL default nomic-embed-text
* ENGRAM_SEMANTIC_MIN cosine threshold for a pure-semantic match (def 0.6)
* */
static double engram_semantic_min(void) {
static double v = -1.0;
if (v >= 0.0) return v;
const char* s = getenv("ENGRAM_SEMANTIC_MIN");
double d = 0.6;
if (s && *s) { char* e = NULL; double t = strtod(s, &e);
if (e != s && t >= 0.0 && t <= 1.0) d = t; }
v = d; return v;
}
#ifdef HAVE_CURL
typedef struct { char* id; uint64_t hash; float* vec; int dim; } EngramEmbEntry;
static EngramEmbEntry* g_emb_items = NULL;
static int64_t g_emb_count = 0, g_emb_cap = 0;
static int g_emb_state = 0; /* 0=unprobed, 1=available, -1=disabled */
static uint64_t engram_fnv1a(const char* s) {
uint64_t h = 1469598103934665603ULL;
if (s) for (const unsigned char* p = (const unsigned char*)s; *p; p++) {
h ^= *p; h *= 1099511628211ULL;
}
return h;
}
/* Parse "embedding":[f,f,...] from an Ollama response. malloc'd vec, or NULL. */
static float* engram_parse_embedding(const char* json, int* out_dim) {
if (!json) return NULL;
const char* p = strstr(json, "\"embedding\"");
if (!p) return NULL;
p = strchr(p, '[');
if (!p) return NULL;
p++;
int cap = 1024, n = 0;
float* v = malloc((size_t)cap * sizeof(float));
if (!v) return NULL;
while (*p && *p != ']') {
while (*p == ' ' || *p == '\t' || *p == '\n' || *p == '\r' || *p == ',') p++;
if (*p == ']' || !*p) break;
char* e = NULL;
double d = strtod(p, &e);
if (e == p) break;
if (n >= cap) { cap *= 2; float* nv = realloc(v, (size_t)cap * sizeof(float));
if (!nv) { free(v); return NULL; } v = nv; }
v[n++] = (float)d;
p = e;
}
if (n == 0) { free(v); return NULL; }
*out_dim = n;
return v;
}
/* JSON-escape src into a malloc'd buffer (no surrounding quotes). */
static char* engram_json_escape(const char* src) {
if (!src) src = "";
size_t n = strlen(src);
char* out = malloc(n * 2 + 1);
if (!out) return NULL;
size_t j = 0;
for (size_t i = 0; i < n; i++) {
unsigned char c = (unsigned char)src[i];
if (c == '"') { out[j++] = '\\'; out[j++] = '"'; }
else if (c == '\\') { out[j++] = '\\'; out[j++] = '\\'; }
else if (c == '\n') { out[j++] = '\\'; out[j++] = 'n'; }
else if (c == '\r') { out[j++] = '\\'; out[j++] = 'r'; }
else if (c == '\t') { out[j++] = '\\'; out[j++] = 't'; }
else if (c < 0x20) { /* drop other control bytes */ }
else { out[j++] = (char)c; }
}
out[j] = '\0';
return out;
}
/* Embed `prefix+text` via Ollama. Returns malloc'd vec (caller frees), or NULL. */
static float* engram_embed_raw(const char* prefix, const char* text, int* out_dim) {
if (!text) return NULL;
const char* url = getenv("ENGRAM_EMBED_URL");
if (!url || !*url) url = "http://localhost:11434/api/embeddings";
const char* model = getenv("ENGRAM_EMBED_MODEL");
if (!model || !*model) model = "nomic-embed-text";
/* Bound content length to keep latency/memory sane on huge nodes. */
char* trunc = NULL;
size_t maxlen = 8192;
if (strlen(text) > maxlen) {
trunc = malloc(maxlen + 1);
if (trunc) { memcpy(trunc, text, maxlen); trunc[maxlen] = '\0'; text = trunc; }
}
char* esc_prefix = engram_json_escape(prefix ? prefix : "");
char* esc = engram_json_escape(text);
free(trunc);
if (!esc || !esc_prefix) { free(esc); free(esc_prefix); return NULL; }
size_t blen = strlen(esc) + strlen(esc_prefix) + strlen(model) + 64;
char* body = malloc(blen);
if (!body) { free(esc); free(esc_prefix); return NULL; }
snprintf(body, blen, "{\"model\":\"%s\",\"prompt\":\"%s%s\"}", model, esc_prefix, esc);
free(esc); free(esc_prefix);
CURL* c = curl_easy_init();
if (!c) { free(body); return NULL; }
HttpBuf rb; httpbuf_init(&rb);
struct curl_slist* h = curl_slist_append(NULL, "Content-Type: application/json");
char errbuf[CURL_ERROR_SIZE]; errbuf[0] = '\0';
curl_easy_setopt(c, CURLOPT_URL, url);
curl_easy_setopt(c, CURLOPT_WRITEFUNCTION, http_write_cb);
curl_easy_setopt(c, CURLOPT_WRITEDATA, &rb);
curl_easy_setopt(c, CURLOPT_POST, 1L);
curl_easy_setopt(c, CURLOPT_POSTFIELDS, body);
curl_easy_setopt(c, CURLOPT_POSTFIELDSIZE, (long)strlen(body));
curl_easy_setopt(c, CURLOPT_HTTPHEADER, h);
curl_easy_setopt(c, CURLOPT_TIMEOUT_MS, el_http_timeout_ms());
curl_easy_setopt(c, CURLOPT_NOSIGNAL, 1L);
curl_easy_setopt(c, CURLOPT_ERRORBUFFER, errbuf);
CURLcode rc = curl_easy_perform(c);
curl_slist_free_all(h);
curl_easy_cleanup(c);
free(body);
if (rc != CURLE_OK) { free(rb.data); return NULL; }
float* v = engram_parse_embedding(rb.data, out_dim);
free(rb.data);
return v;
}
/* One-shot probe: is semantic search available? Caches the verdict. */
static int engram_semantic_enabled(void) {
if (g_emb_state != 0) return g_emb_state == 1;
const char* s = getenv("ENGRAM_SEMANTIC");
if (s && strcmp(s, "0") == 0) { g_emb_state = -1; return 0; }
int dim = 0;
float* v = engram_embed_raw("search_query: ", "probe", &dim);
if (v && dim > 0) { free(v); g_emb_state = 1; return 1; }
free(v);
g_emb_state = -1; return 0;
}
/* Embed the query. Returns malloc'd vec (caller frees), or NULL if semantic off. */
static float* engram_embed_query(const char* q, int* dim) {
if (!engram_semantic_enabled()) return NULL;
if (!q || !*q) return NULL;
return engram_embed_raw("search_query: ", q, dim);
}
/* Cached node embedding. Returns a pointer OWNED BY THE CACHE — do not free. */
static const float* engram_node_vec(EngramNode* n, int* out_dim) {
if (!n || !n->id) return NULL;
uint64_t h = engram_fnv1a(n->content);
for (int64_t i = 0; i < g_emb_count; i++) {
if (g_emb_items[i].id && strcmp(g_emb_items[i].id, n->id) == 0) {
if (g_emb_items[i].hash == h && g_emb_items[i].vec) {
*out_dim = g_emb_items[i].dim; return g_emb_items[i].vec;
}
/* content changed → re-embed in place */
int dim = 0;
float* v = engram_embed_raw("search_document: ", n->content ? n->content : "", &dim);
if (!v) return NULL;
free(g_emb_items[i].vec);
g_emb_items[i].vec = v; g_emb_items[i].dim = dim; g_emb_items[i].hash = h;
*out_dim = dim; return v;
}
}
int dim = 0;
float* v = engram_embed_raw("search_document: ", n->content ? n->content : "", &dim);
if (!v) return NULL;
if (g_emb_count >= g_emb_cap) {
int64_t nc = g_emb_cap ? g_emb_cap * 2 : 256;
EngramEmbEntry* ni = realloc(g_emb_items, (size_t)nc * sizeof(EngramEmbEntry));
if (!ni) { free(v); return NULL; }
g_emb_items = ni; g_emb_cap = nc;
}
g_emb_items[g_emb_count].id = strdup(n->id);
g_emb_items[g_emb_count].hash = h;
g_emb_items[g_emb_count].vec = v;
g_emb_items[g_emb_count].dim = dim;
g_emb_count++;
*out_dim = dim; return v;
}
static double engram_cosine(const float* a, const float* b, int dim) {
double dot = 0, na = 0, nb = 0;
for (int i = 0; i < dim; i++) { dot += (double)a[i] * b[i];
na += (double)a[i] * a[i];
nb += (double)b[i] * b[i]; }
if (na <= 0 || nb <= 0) return 0.0;
return dot / (sqrt(na) * sqrt(nb));
}
/* Cosine of node n against the query vector; 0 if unavailable / dim mismatch. */
static double engram_node_cosine(EngramNode* n, const float* qvec, int qdim) {
if (!qvec || qdim <= 0) return 0.0;
int ndim = 0;
const float* nv = engram_node_vec(n, &ndim);
if (!nv || ndim != qdim) return 0.0;
return engram_cosine(qvec, nv, qdim);
}
#else /* !HAVE_CURL — semantic layer compiled out; callers stay pure-lexical.
* Only the two boundary functions the always-compiled search/activate
* code calls are stubbed; the query embed always yields NULL so every
* cosine is 0 and every caller collapses to lexical-only. */
static float* engram_embed_query(const char* q, int* dim) { (void)q; (void)dim; return NULL; }
static double engram_node_cosine(EngramNode* n, const float* qvec, int qdim) {
(void)n; (void)qvec; (void)qdim; return 0.0;
}
#endif /* HAVE_CURL */
el_val_t engram_search(el_val_t query, el_val_t limit) {
EngramStore* g = engram_get();
const char* q = EL_CSTR(query);
@@ -7092,8 +7179,15 @@ el_val_t engram_search(el_val_t query, el_val_t limit) {
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
int ntok = engram_tokenize_query(q, toks, ENGRAM_MAX_QTOKENS);
if (ntok == 0) return lst;
/* Semantic augmentation: embed the query once; a node is a hit if it covers
* >=1 query token (tokenized-lexical, #66) OR its cosine clears the
* threshold (#67). qvec is NULL (cosine 0) when semantic is unavailable
* pure tokenized-lexical, byte-identical to the lexical-only behaviour. */
int qdim = 0;
float* qvec = engram_embed_query(q, &qdim);
double sem_min = engram_semantic_min();
EngramRankEntry* hits = malloc((size_t)g->node_count * sizeof(EngramRankEntry));
if (!hits) return lst;
if (!hits) { free(qvec); return lst; }
int64_t nhits = 0;
for (int64_t i = 0; i < g->node_count; i++) {
EngramNode* n = &g->nodes[i];
@@ -7103,20 +7197,24 @@ el_val_t engram_search(el_val_t query, el_val_t limit) {
* + engram_compile_layered_json that's the legitimate path. */
if (engram_layer_is_transparent(n->layer_id)) continue;
int sc = engram_node_match_score(n, toks, ntok);
if (sc > 0) {
double sem = qvec ? engram_node_cosine(n, qvec, qdim) : 0.0;
if (sc > 0 || sem >= sem_min) {
hits[nhits].idx = i;
hits[nhits].score = sc;
hits[nhits].salience = n->salience;
nhits++;
}
}
/* Rank by distinct tokens matched (desc) then salience (desc), then cap. */
/* Rank by distinct tokens matched (desc) then salience (desc), then cap.
* Pure-semantic hits (token score 0) sort after every lexical hit a
* lexical semantic union with lexical precedence. */
qsort(hits, (size_t)nhits, sizeof(EngramRankEntry), engram_rank_cmp);
int64_t end = nhits < lim ? nhits : lim;
for (int64_t k = 0; k < end; k++) {
lst = el_list_append(lst, engram_node_to_map(&g->nodes[hits[k].idx]));
}
free(hits);
free(qvec);
return lst;
}
@@ -7439,48 +7537,6 @@ static double engram_goal_bias(const EngramNode* n, const char* query) {
return bias;
}
/* eg_cmp_double_desc — qsort comparator, descending doubles. */
static int eg_cmp_double_desc(const void* a, const void* b) {
double da = *(const double*)a, db = *(const double*)b;
if (da < db) return 1;
if (da > db) return -1;
return 0;
}
/* eg_enforce_wm_cap_global — clamp the store-wide working-memory population
* to ENGRAM_WM_CAP, keeping the top-K by current weight. Runs at every point
* that materializes WM: post-activation persist and snapshot load/merge.
* (Ported from release runtime v1.0.0-20260501 Pass 5, 2026-07-15.) */
static void eg_enforce_wm_cap_global(EngramStore* g) {
int64_t wm_count = 0;
for (int64_t i = 0; i < g->node_count; i++) {
if (g->nodes[i].working_memory_weight > 0.0) wm_count++;
}
if (wm_count <= ENGRAM_WM_CAP) return;
double* vals = malloc((size_t)wm_count * sizeof(double));
if (!vals) return; /* OOM: over cap this call, no corruption */
int64_t vi = 0;
for (int64_t i = 0; i < g->node_count; i++) {
if (g->nodes[i].working_memory_weight > 0.0)
vals[vi++] = g->nodes[i].working_memory_weight;
}
qsort(vals, (size_t)wm_count, sizeof(double), eg_cmp_double_desc);
double cutoff = vals[ENGRAM_WM_CAP - 1];
free(vals);
int64_t above = 0;
for (int64_t i = 0; i < g->node_count; i++) {
if (g->nodes[i].working_memory_weight > cutoff) above++;
}
int64_t slots_at_cutoff = ENGRAM_WM_CAP - above;
for (int64_t i = 0; i < g->node_count; i++) {
EngramNode* n = &g->nodes[i];
if (n->working_memory_weight <= 0.0) continue;
if (n->working_memory_weight > cutoff) continue;
if (slots_at_cutoff > 0) { slots_at_cutoff--; continue; }
n->working_memory_weight = 0.0; /* evict: over global cap */
}
}
el_val_t engram_activate(el_val_t query, el_val_t depth) {
EngramStore* g = engram_get();
const char* q = EL_CSTR(query);
@@ -7508,21 +7564,31 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
if (!seeds) {
free(best_bg); free(best_hops); free(reached); return out;
}
/* Tokenize once: a node seeds if it matches ANY query token, and its seed
* activation is scaled by token coverage (fraction of distinct query
* tokens it contains) so a node matching all words seeds more strongly
* than one matching a single word. Single-word queries coverage 1.0,
* identical to the prior whole-query behavior. */
/* Tokenized + semantic seeding: a node seeds if it covers >=1 query token
* (tokenized-lexical, #66) OR its cosine clears the threshold (#67). A
* lexical seed's activation is scaled by token coverage (fraction of
* distinct query tokens covered) so a node matching all words seeds more
* strongly than one matching a single word; single-word queries coverage
* 1.0. A pure-semantic seed (no token match) is instead down-weighted by
* its cosine so paraphrase matches spread without overpowering exact seeds.
* q_vec is NULL (cosine 0) when semantic is unavailable the seed set is
* exactly the tokenized-lexical one. q_vec is freed right after this loop
* so the many downstream early-returns need no cleanup change. */
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
int ntok = engram_tokenize_query(q, toks, ENGRAM_MAX_QTOKENS);
int q_dim = 0;
float* q_vec = engram_embed_query(q, &q_dim);
double q_sem_min = engram_semantic_min();
for (int64_t i = 0; i < g->node_count; i++) {
EngramNode* n = &g->nodes[i];
int sc = engram_node_match_score(n, toks, ntok);
if (sc > 0) {
double sem = q_vec ? engram_node_cosine(n, q_vec, q_dim) : 0.0;
if (sc > 0 || sem >= q_sem_min) {
double tdecay = engram_temporal_decay(n, now_ms);
double dampen = engram_activation_dampen(n);
double cover = ntok > 0 ? (double)sc / (double)ntok : 1.0;
double act = n->salience * tdecay * dampen * cover;
double act = n->salience * tdecay * dampen;
if (sc > 0) act *= (ntok > 0 ? (double)sc / (double)ntok : 1.0);
else act *= sem; /* pure-semantic seed: down-weight by cosine */
seeds[seed_count].idx = i;
seeds[seed_count].act = act;
seeds[seed_count].created_at = n->created_at;
@@ -7532,6 +7598,7 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
reached[i] = 1;
}
}
free(q_vec);
/* Compute mean seed created_at for temporal proximity bonus. */
int64_t seed_epoch = 0;
if (seed_count > 0) {
@@ -7686,12 +7753,6 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
g->nodes[i].working_memory_weight = wm_weights[i];
}
/* Global WM cap: keep only the top ENGRAM_WM_CAP by weight across the
* whole store (see eg_enforce_wm_cap_global). Without this, repeated
* activation calls accumulate hundreds of "promoted" nodes and WM stops
* meaning anything (production heartbeats showed wm_active up to 778). */
eg_enforce_wm_cap_global(g);
/* ── Collect all background-activated nodes for the return value ────
* Callers see both layers. Context compilation uses only promoted nodes
* (working_memory_weight > 0). Sort: promoted first by wm_weight desc,
@@ -8069,9 +8130,6 @@ el_val_t engram_load(el_val_t path) {
}
}
}
/* WM cap discipline applies to every entry point that materializes WM,
* including snapshot restore (see eg_enforce_wm_cap_global). */
eg_enforce_wm_cap_global(g);
free(data);
return 1;
}
@@ -8096,14 +8154,16 @@ el_val_t engram_get_node_json(el_val_t id) {
* matches the given string. Returns the node as a JSON object string, or "{}"
* if no match is found.
*
* Exact match (strcmp, not substring) because labels like "conv:history"
* must not collide with nodes whose content contains that substring.
* Used by chat.el to retrieve well-known nodes (e.g. "conv:history",
* "session:summary") by their stable label rather than by ID, which is immune
* to vector index drift across restarts.
*
* Ported from the release runtime 2026-07-16 self-review: chat.el has called
* this since 2026-07-01 but the function only existed in
* releases/v1.0.0-20260501/el_runtime.c the soul daemon (which builds
* against THIS runtime) failed to compile once clang made implicit
* declarations an error. */
* Exact match (strcmp, not istr_contains) because labels like "conv:history"
* must not collide with nodes whose content happens to contain that substring.
*
* Backported verbatim (idiom-adapted to jb_finish) from release runtime
* v1.0.0-20260501 to unblock the soul regen link: chat.el references this
* native but the current runtime lacked its definition. */
el_val_t engram_get_node_by_label(el_val_t label) {
const char* lbl = EL_CSTR(label);
if (!lbl || !*lbl) return el_wrap_str(el_strdup("{}"));
@@ -8126,39 +8186,50 @@ el_val_t engram_search_json(el_val_t query, el_val_t limit) {
if (lim <= 0) lim = 100;
JsonBuf b; jb_init(&b);
jb_putc(&b, '[');
int first = 1;
if (q && *q) {
if (q && *q && g->node_count > 0) {
/* Collect candidates from the UNION of tokenized-lexical and semantic
* matches, score each, rank by score, emit the top `lim`. A node is a
* candidate if it covers >=1 query token (tokenized-lexical, #66) OR its
* query cosine clears the threshold (#67). Lexical score is the distinct
* token count (>=1), so any lexical hit outranks a pure-semantic hit
* (cosine < 1); pure-semantic hits are scored by cosine alone. When
* semantic is unavailable qvec is NULL, sem is 0, only tokenized-lexical
* hits are collected, and the stable insertion sort preserves order. */
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
int ntok = engram_tokenize_query(q, toks, ENGRAM_MAX_QTOKENS);
if (ntok > 0) {
EngramRankEntry* hits =
malloc((size_t)g->node_count * sizeof(EngramRankEntry));
if (hits) {
int64_t nhits = 0;
for (int64_t i = 0; i < g->node_count; i++) {
EngramNode* n = &g->nodes[i];
/* Filter transparent layers — same as engram_search. */
if (engram_layer_is_transparent(n->layer_id)) continue;
int sc = engram_node_match_score(n, toks, ntok);
if (sc > 0) {
hits[nhits].idx = i;
hits[nhits].score = sc;
hits[nhits].salience = n->salience;
nhits++;
}
int qdim = 0;
float* qvec = engram_embed_query(q, &qdim);
double sem_min = engram_semantic_min();
typedef struct { int64_t idx; double score; } Cand;
Cand* cand = malloc((size_t)g->node_count * sizeof(Cand));
if (cand) {
int64_t nc = 0;
for (int64_t i = 0; i < g->node_count; i++) {
EngramNode* n = &g->nodes[i];
if (engram_layer_is_transparent(n->layer_id)) continue;
int sc = engram_node_match_score(n, toks, ntok);
double sem = qvec ? engram_node_cosine(n, qvec, qdim) : 0.0;
if (sc > 0 || sem >= sem_min) {
cand[nc].idx = i;
cand[nc].score = (double)sc + sem;
nc++;
}
/* Rank by distinct tokens matched (desc) then salience (desc). */
qsort(hits, (size_t)nhits, sizeof(EngramRankEntry),
engram_rank_cmp);
int64_t end = nhits < lim ? nhits : lim;
for (int64_t k = 0; k < end; k++) {
if (!first) jb_putc(&b, ',');
engram_emit_node_json(&b, &g->nodes[hits[k].idx]);
first = 0;
}
free(hits);
}
/* Insertion sort by score desc; stable for equal scores. */
for (int64_t i = 1; i < nc; i++) {
Cand k = cand[i]; int64_t j = i - 1;
while (j >= 0 && cand[j].score < k.score) { cand[j + 1] = cand[j]; j--; }
cand[j + 1] = k;
}
int first = 1;
for (int64_t i = 0; i < nc && i < lim; i++) {
if (!first) jb_putc(&b, ',');
engram_emit_node_json(&b, &g->nodes[cand[i].idx]);
first = 0;
}
free(cand);
}
free(qvec);
}
jb_putc(&b, ']');
return el_wrap_str(jb_finish(&b));
@@ -8676,9 +8747,6 @@ el_val_t engram_load_merge(el_val_t path) {
}
}
/* Merged nodes can carry snapshot WM weights too — hold the cap here as
* well (see eg_enforce_wm_cap_global). */
eg_enforce_wm_cap_global(g);
free(data);
return (el_val_t)added_nodes;
}
+1
View File
@@ -1072,6 +1072,7 @@ el_val_t __engram_save(el_val_t path) { return engram_save
el_val_t __engram_load(el_val_t path) { return engram_load(path); }
el_val_t __engram_get_node_json(el_val_t id) { return engram_get_node_json(id); }
el_val_t __engram_get_node_by_label(el_val_t label) { return engram_get_node_by_label(label); }
el_val_t __engram_search_json(el_val_t query, el_val_t limit) {
return engram_search_json(query, limit);
+1
View File
@@ -226,6 +226,7 @@ el_val_t __engram_activate(el_val_t query, el_val_t depth);
el_val_t __engram_save(el_val_t path);
el_val_t __engram_load(el_val_t path);
el_val_t __engram_get_node_json(el_val_t id);
el_val_t __engram_get_node_by_label(el_val_t label);
el_val_t __engram_search_json(el_val_t query, el_val_t limit);
el_val_t __engram_scan_nodes_json(el_val_t limit, el_val_t offset);
el_val_t __engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, el_val_t offset);
+1
View File
@@ -2670,6 +2670,7 @@ fn builtin_arity(name: String) -> Int {
if str_eq(name, "engram_save") { return 1 }
if str_eq(name, "engram_load") { return 1 }
if str_eq(name, "engram_get_node_json") { return 1 }
if str_eq(name, "engram_get_node_by_label") { return 1 }
if str_eq(name, "engram_search_json") { return 2 }
if str_eq(name, "engram_scan_nodes_json") { return 2 }
if str_eq(name, "engram_neighbors_json") { return 3 }
@@ -0,0 +1,186 @@
#ifndef EL_PLATFORM_WIN_H
#define EL_PLATFORM_WIN_H
/*
* el_platform_win.h Windows OS-boundary shim for el_runtime.c.
*
* Branch: feat/windows-el-runtime. Included ONLY when _WIN32 is defined; the POSIX build is
* untouched. Goal: let el_runtime.c (a BSD-sockets / dlfcn / fork host) compile and link with
* mingw-w64 into a native neuron.exe, with no behavioural change to the Linux/macOS build.
*
* What it maps:
* - sockets : winsock2 (same call names: socket/bind/listen/accept/recv/send/setsockopt).
* Sockets close with closesocket() (see el_closesocket), and the stack must be
* started once with WSAStartup done automatically via a load-time constructor.
* - dlsym : el_runtime.c uses dlsym(RTLD_DEFAULT, name) to resolve callback/tool symbols
* exported by the main module. Windows equivalent: GetProcAddress on the process
* module. Link the soul with -Wl,--export-all-symbols so the symbols are findable.
* - popen : mapped to _popen/_pclose.
* - threads : UNCHANGED. mingw-w64 ships winpthreads, so <pthread.h> + -lpthread just work.
*/
#ifndef WIN32_LEAN_AND_MEAN
#define WIN32_LEAN_AND_MEAN
#endif
#include <winsock2.h>
#include <ws2tcpip.h>
#include <windows.h>
#include <io.h>
#include <process.h>
/* Portable headers mingw-w64 provides (verified present). */
#include <stdarg.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <strings.h> /* strcasecmp */
#include <ctype.h>
#include <math.h>
#include <time.h>
#include <sys/time.h> /* mingw-w64 provides gettimeofday here */
#include <sys/types.h>
#include <sys/stat.h>
#include <fcntl.h>
#include <dirent.h>
#include <errno.h>
#include <pthread.h>
/* ── socket close ─────────────────────────────────────────────────────────── */
/* Winsock closes sockets with closesocket(), not close() (close() is for file fds). The POSIX
build defines the same helper as close() so the call sites are identical across platforms. */
static inline int el_closesocket(SOCKET s) { return closesocket(s); }
/* ── setsockopt optval type ───────────────────────────────────────────────── */
/* Winsock's setsockopt takes optval as (const char*); POSIX takes (const void*), so el_runtime.c
passes &int directly. GCC 14+ makes that an error under -Wincompatible-pointer-types. Wrap it so
the runtime's POSIX-style call sites compile unchanged (defined before the macro so the wrapper
itself resolves to the real winsock setsockopt). */
static inline int el_setsockopt(SOCKET s, int level, int optname, const void* optval, int optlen) {
return setsockopt(s, level, optname, (const char*)optval, optlen);
}
#define setsockopt(s, l, o, v, n) el_setsockopt((s), (l), (o), (v), (int)(n))
/* ── winsock init (once, at load) ─────────────────────────────────────────── */
static void el__win_net_init(void) {
static int inited = 0;
if (!inited) { WSADATA w; WSAStartup(MAKEWORD(2, 2), &w); inited = 1; }
}
__attribute__((constructor)) static void el__win_ctor(void) { el__win_net_init(); }
/* ── dlsym → GetProcAddress ───────────────────────────────────────────────── */
#ifndef RTLD_DEFAULT
#define RTLD_DEFAULT ((void*)0)
#endif
static inline void* el_win_dlsym(void* handle, const char* name) {
(void)handle;
return (void*)(uintptr_t)GetProcAddress(GetModuleHandleA(NULL), name);
}
#define dlsym(h, n) el_win_dlsym((h), (n))
/* ── popen / pclose ───────────────────────────────────────────────────────── */
#define popen _popen
#define pclose _pclose
/* ── misc POSIX → Win32 shims ─────────────────────────────────────────────── */
#include <direct.h> /* _mkdir */
#define mkdir(path, mode) _mkdir(path) /* POSIX mkdir(path,mode) → _mkdir(path) */
#define timegm _mkgmtime /* UTC tm → time_t */
/* setenv/unsetenv: not in the Windows CRT; map to _putenv_s / SetEnvironmentVariable. */
static inline int setenv(const char* name, const char* value, int overwrite) {
(void)overwrite;
return _putenv_s(name, value ? value : "");
}
static inline int unsetenv(const char* name) {
/* _putenv_s(name, "") sets VAR="" rather than removing it.
* SetEnvironmentVariableA(name, NULL) truly deletes it from the Win32
* env block; then we sync the CRT cache with _putenv("NAME="). */
SetEnvironmentVariableA(name, NULL);
size_t len = strlen(name);
char *buf = (char*)malloc(len + 2);
if (!buf) return -1;
memcpy(buf, name, len);
buf[len] = '=';
buf[len + 1] = '\0';
_putenv(buf);
free(buf);
return 0;
}
/* nanosleep — not available in MSVC/UCRT; approximate with Sleep(). */
static inline int el_nanosleep(const struct timespec *req, struct timespec *rem) {
(void)rem;
DWORD ms = (DWORD)((req->tv_sec * 1000ULL) + (req->tv_nsec / 1000000ULL));
Sleep(ms ? ms : 1);
return 0;
}
#define nanosleep(req, rem) el_nanosleep((req), (rem))
/* localtime_r/gmtime_r: Windows offers localtime_s/gmtime_s with reversed arg order. */
static inline struct tm* localtime_r(const time_t* t, struct tm* out) {
return localtime_s(out, t) == 0 ? out : (struct tm*)0;
}
static inline struct tm* gmtime_r(const time_t* t, struct tm* out) {
return gmtime_s(out, t) == 0 ? out : (struct tm*)0;
}
/* ── libcurl: degradable stubs for the curl-less Windows build ─────────────── */
/* The curl-less validation build (WITH_CURL=0) links no libcurl. el_runtime.c uses libcurl
* unconditionally for its HTTP client / LLM layer; these stubs let it compile and link so the
* runtime, HTTP *server*, graph and memory work natively on Windows. Live outbound HTTP/LLM calls
* degrade to a runtime error (curl_easy_perform returns an error) matching the documented
* curl-less contract. When HAVE_CURL is defined (WITH_CURL=1) the real <curl/curl.h> is used and
* this whole block is compiled out. POSIX never sees this header, so the POSIX build is untouched. */
#ifndef HAVE_CURL
typedef void CURL;
typedef int CURLcode;
#define CURLE_OK 0
#define CURLE_HTTP_RETURNED_ERROR 22
#define CURL_ERROR_SIZE 256
/* Option ids: values are irrelevant to the no-op setopt below; kept distinct for readability. */
#define CURLOPT_URL 10002
#define CURLOPT_WRITEFUNCTION 20011
#define CURLOPT_WRITEDATA 10001
#define CURLOPT_POSTFIELDS 10015
#define CURLOPT_POSTFIELDSIZE 120
#define CURLOPT_POST 47
#define CURLOPT_HTTPHEADER 10023
#define CURLOPT_TIMEOUT_MS 155
#define CURLOPT_NOSIGNAL 99
#define CURLOPT_USERAGENT 10018
#define CURLOPT_FOLLOWLOCATION 52
#define CURLOPT_ERRORBUFFER 10010
#define CURLOPT_CUSTOMREQUEST 10036
#define CURLOPT_FAILONERROR 45
struct curl_slist { char* data; struct curl_slist* next; };
static inline struct curl_slist* curl_slist_append(struct curl_slist* list, const char* s) {
struct curl_slist* node = (struct curl_slist*)malloc(sizeof(struct curl_slist));
if (!node) return list;
node->data = s ? strdup(s) : NULL;
node->next = NULL;
if (!list) return node;
struct curl_slist* p = list;
while (p->next) p = p->next;
p->next = node;
return list;
}
static inline void curl_slist_free_all(struct curl_slist* list) {
while (list) { struct curl_slist* n = list->next; free(list->data); free(list); list = n; }
}
static inline CURL* curl_easy_init(void) { return (CURL*)malloc(1); }
static inline CURLcode curl_easy_setopt(CURL* h, int opt, ...) { (void)h; (void)opt; return CURLE_OK; }
static inline CURLcode curl_easy_perform(CURL* h) { (void)h; return 7 /* CURLE_COULDNT_CONNECT */; }
static inline void curl_easy_cleanup(CURL* h) { free(h); }
static inline const char* curl_easy_strerror(CURLcode c) {
(void)c; return "libcurl not built in (curl-less build)";
}
#endif /* !HAVE_CURL */
#endif /* EL_PLATFORM_WIN_H */
File diff suppressed because it is too large Load Diff
+12 -36
View File
@@ -117,15 +117,6 @@ el_val_t el_min(el_val_t a, el_val_t b);
void el_retain(el_val_t v);
void el_release(el_val_t v);
/* ── Arena scoping ────────────────────────────────────────────────────────────
* el_arena_push() activates the string arena (if not already active) and
* returns a mark; el_arena_pop(mark) frees all strings allocated since that
* mark. Used by codegen for per-function/statement scoping and by long-running
* EL loops (e.g. the soul daemon's awareness tick) to reclaim per-iteration
* allocations. */
el_val_t el_arena_push(void);
el_val_t el_arena_pop(el_val_t mark);
/* ── List ────────────────────────────────────────────────────────────────── */
el_val_t el_list_new(el_val_t count, ...);
@@ -151,7 +142,6 @@ el_val_t http_get_with_headers(el_val_t url, el_val_t headers_map);
el_val_t http_post_with_headers(el_val_t url, el_val_t body, el_val_t headers_map);
el_val_t http_post_form_auth(el_val_t url, el_val_t form_body, el_val_t auth_header);
el_val_t http_delete(el_val_t url);
el_val_t http_delete_json(el_val_t url, el_val_t json_body);
void http_serve(el_val_t port, el_val_t handler);
void http_set_handler(el_val_t name);
@@ -177,11 +167,6 @@ void http_set_handler(el_val_t name);
void http_serve_v2(el_val_t port, el_val_t handler);
void http_set_handler_v2(el_val_t name);
/* Non-blocking variant of http_serve: runs the accept loop in a background
* pthread and returns immediately so the caller can continue (used by the
* soul daemon to run awareness_run() after starting its HTTP API). */
void http_serve_async(el_val_t port, el_val_t handler);
/* Build an HTTP response envelope. `headers_json` should be a JSON object
* literal like `{"WWW-Authenticate":"Basic"}` (or "" / "{}" for none). The
* returned string carries the discriminator `{"el_http_response":1,...}`
@@ -591,7 +576,6 @@ el_val_t engram_list_layers(void);
el_val_t engram_get_node(el_val_t id);
void engram_strengthen(el_val_t node_id);
void engram_forget(el_val_t node_id);
el_val_t engram_prune_telemetry(el_val_t older_than_ms);
el_val_t engram_node_count(void);
el_val_t engram_search(el_val_t query, el_val_t limit);
el_val_t engram_scan_nodes(el_val_t limit, el_val_t offset);
@@ -610,32 +594,12 @@ el_val_t engram_load(el_val_t path);
* can pass results straight through without round-tripping ElList/ElMap
* through json_stringify. */
el_val_t engram_get_node_json(el_val_t id);
el_val_t engram_get_node_by_label(el_val_t label);
el_val_t engram_search_json(el_val_t query, el_val_t limit);
el_val_t engram_scan_nodes_json(el_val_t limit, el_val_t offset);
el_val_t engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, el_val_t offset);
el_val_t engram_neighbors_json(el_val_t node_id, el_val_t max_depth, el_val_t direction);
el_val_t engram_activate_json(el_val_t query, el_val_t depth);
el_val_t engram_stats_json(void);
el_val_t engram_act_stats_json(void);
el_val_t engram_text_health_json(void);
el_val_t engram_cosine_sim(el_val_t id_a, el_val_t id_b);
/* Destructively pop up to `max` newly-formed Hebbian associations as a JSON
* array of {from_id,to_id,weight,hebb}. The learning process (soul daemon) is
* not the process that owns persistence (engram HTTP server); this is how a
* self-formed association crosses that boundary. (2026-08-07 self-review.) */
el_val_t engram_hebb_drain_json(el_val_t max);
/* Document frequency of a term across node labels — term-specificity signal
* for curiosity seed selection. (2026-08-03 self-review.) */
el_val_t engram_label_df(el_val_t term);
/* Best curiosity seed from one node: argmax over idf·position·casing across
* the candidate tokens of its label, falling back to its content when the
* label is a sentinel. Excludes pipe-delimited tabu terms during selection
* and gates candidates to the df band [min_df, max_df]. Returns "" when
* nothing qualifies. (2026-08-13 self-review.) */
el_val_t engram_salient_term(el_val_t node_id, el_val_t max_df,
el_val_t min_df, el_val_t tabu);
el_val_t engram_embed_backfill(el_val_t count);
el_val_t engram_list_layers_json(void);
/* Working memory introspection — count, mean weight, and top-N snapshot.
* Ported from el-compiler/runtime on 2026-06-30 self-review. */
@@ -794,6 +758,18 @@ el_val_t trace_span_start(el_val_t name);
el_val_t trace_span_end(el_val_t span_handle);
el_val_t emit_event(el_val_t name, el_val_t duration_ms);
/* ── Runtime symbols required by the soul modules ──────────────────────────── */
/* All implemented in el_runtime.c but omitted from this release header; the soul dist modules
* reference them directly, so the public header must export them. Declarations only mirrors the
* mainline el_runtime.h and is platform-independent (no behavioural change to the POSIX build). */
typedef el_val_t (*http_handler_fn)(el_val_t method, el_val_t path, el_val_t body);
typedef el_val_t (*http_handler4_fn)(el_val_t method, el_val_t path, el_val_t body, el_val_t headers);
el_val_t el_arena_push(void);
el_val_t el_arena_pop(el_val_t mark);
void http_serve_async(el_val_t port, el_val_t handler);
el_val_t engram_get_node_by_label(el_val_t label);
el_val_t engram_prune_telemetry(el_val_t older_than_ms);
#ifdef __cplusplus
}
#endif