Rules for future sessions in this repo, captured after corrections and mistakes.
- All temp/working files go inside the morpheus folder (
tmp/, git-ignored) — never in~/.lmstudio/scratchpads. (User correction, 2026-10-07.)
- The stemlib noun sources are named
nom01…nom36(no dot): the globstemsrc/*.nom*silently matches onlylsj.nom,nom.biblical, … Usestemsrc/nom*when searching all of them. - zsh chokes on unquoted parentheses in command lists; quote patterns or use a
script file for anything with
( ). - When scrubbing a name from the repo, grep case-insensitively with word
boundaries over ALL file types — not just
.py/.md/.sh. Build artifacts (e.g.stemlib/Greek/conjfile, generated by concatenating stemsrc), generated reports, and saved snapshots all retain old strings long after the sources are clean. Fix at the root: edit the source, regenerate the artifact, delete stale snapshots.
- cruncher only matches lowercase beta code; uppercase input (reference-
corpus style) silently returns nothing. The toolkit now normalizes this
(
runner._normalize_beta); keep it that way and test both cases. to_beta()drops a breathing when the Unicode letter has none: plain Ι gives*israhl, but Ἰ (U+1F08) gives*)israh/l. When verifying name entries, test with properly breathed/accented forms AND check-n(accent-insensitive) mode.- Explicit
:wd:noun/name forms match exactly (diacritics included); the accent-insensitive pass (cruncher -S -n) is what makes unaccented variants hit. - Generative
:no:<stem> os_ou …patterns do not work for every stem shape (worked for σάββατον, produced nothing for τρυβλίον/ἀδελφιδός). Verify each new entry against the built index (steminds/nomind) and fall back to explicit:wd:lines. - The incremental stemlib makefile does not always pick up a brand-new stemsrc
file; if entries are missing from
nomind/vbind, check timestamps and force the target (e.g.touch stemsrc/lsj.nom && make). - Paradigm declarations (
;xx,@ xx) in vbs sources never reach the built index. The build iscat … > conjfile; do_conj; mv conjfile.short /tmp/vbmorph; indexvbs— indexvbs indexes the do_conj output, and only processes:vs/:aj/:no/:vb/:wd/:de:lines (all;,@,--prefixed lines are ignored). So;pp/;pfdeclarations change nothing in vbind; use explicit:vb:lines. - The generative deriv path rejects non-reduplicated perfect stems
(
derivio.c checkcomderiv2:if (!had_redupl && Is_perfect(stemtype)) continue;), so -έω pp participles (εὐλογημένος) never generate. Explicit:vb:lines have no such gate. - GenIrregForm requires a real stemtype in the keys:
o_stemalone is derivtypes-only and fails with "no stemtype seen"; add e.g.omi_prbefore it (precedent: existing:vb:di/dw omi_pr …). FixRecAcc/FixPersAcc preserve any accent already present, so a:vb:/:wd:line storing the corpus's exact surface spelling round-trips exactly through CheckGenWords. standword()rewrites GRAVE→ACUTE (every position) and strips accents after the first on every input before lookup; CheckGenWords compares accents exactly in normal mode (morphstrcmp, only|≡iequivalence) but stripaccs both sides under IGNORE_ACCENTS. Consequence: a stored word-final grave can never match an accented input — store the acute spelling instead (or both spellings, as with αλληλουία). The reference corpus is internally inconsistent (e.g. νοσσιὰν and νοσσιάν for one form; λημφθήτω unaccented in beta but accented in its Greek column) — index every attested standardized spelling.- Noun lookup (
chckindecl) always strips all accents from the key; nomind keys are accent-stripped and merge all full-form variants sharing a stripped key onto one line (morphstrcmp:|≡i, plus "s1 ends ⇒ equal" quirk). The lookup is therefore accent-blind; only CheckGenWords' final compare is accent-sensitive. Multi-entry lines are fine — the parser iterates all entries. - C# port has placeholder stubs (
chckindecl.csreturns 0;CheckGenWordsFuncuses plainstring.Compare, no comptab/accent logic) — don't assume C/C# parity when reasoning about behavior from the.csfiles.
Morpheus.analyze_beta()returnsTokenResultwrappers, not raw analysis lists (result.analyses). The runner-level method returns bare lists; the API level does not.- Reference-corpus parse codes are position-dependent by lexical class: verb codes are 5 chars (tense voice mood person number), nominal codes only 3 (case number gender). Any field-by-field comparison must branch on class AND use per-class minimum lengths, or most nominals silently drop out of the comparison.
- Morpheus's 8-char parsing code contains dashes (
3AAI-S--), so shape-based column detection ("parse = the dash-free [A-Z0-9]+ field") misparses generated files — it reads the parse code as the lemma. Parse generated output positionally. - When training disambiguation models on a gold corpus and then evaluating agreement with that same corpus, say so in the report: the numbers measure domain consistency, not absolute accuracy.
- Never hardcode the reference-corpus location (path or name) in any file.
Scripts take
--corpus-dir/ envLXX_CORPUS_DIR; generated reports use a neutral label. The user greps the working tree — committed files, tmp/ scratch, and build artifacts must all stay clean. - The reference corpus is UPPERCASE beta code with its own dialect quirks
(Round 6): breathings follow the letter (
A)PE..., notAPE...),/is a plain syllable separator inside words (KURI/OU= κυρίου — substring greps likeRIOUmiss it), and\,=,|are accent markers used inconsistently for the same word (αὐτοῦ appears asAU)TOU=,AU)TOU\S,AU)TOU/S). Never assume a spelling: key everything throughcanon()(to_beta∘from_beta+ lowercase), and remember one word can occupy several canon keys that must all be covered by any table. - The eval's "top lemma disagreements" list is misleading (Round 6): it sorts by total form frequency and prints only the first disagreeing line, so a rare quirk on a frequent form looks like a mass error (ὅτι read as ×4,044; actually 7/4044). Always verify with true per-form mismatch counts before sizing a "prize pool".
- Joint-frequency fallback: score unattested (lemma, POS) pairs at a floor of one occurrence — never zero, never full plain frequency (Round 6): the reference's type scheme and Morpheus's POS disagree for the same lemma (αὐτός A vs RD), so zero kills correct candidates; full plain frequency over-credits them (σός(N) 0.432 vs σύ(RP) 0.838 left a gap bigram context flips). An unattested pair can never be the gold reading of any token, so the floor only removes wrong answers.
- The Viterbi tagger keeps one candidate per coarse POS (first on score ties) (Round 6): feature-level ambiguity within a POS — masc/neut homographs, duals — is resolved arbitrarily, and candidates with empty parse codes pass field comparison vacuously (no comparable pairs ⇒ True). Break same-POS ties explicitly (per-spelling reference aggregates; the corpus's accent notation encodes readings) instead of trusting Morpheus's internal candidate order.
Analysis.lemma/.variants()are Unicode, frequency-table keys normalized Unicode (Round 6): compare vianormalize_lemma(), never raw literal equality —from_betacan return NFD and precomposed Greek literals won't match it.
- Augmented assignment on a function call is a SyntaxError:
d.setdefault(k, set()) |= x— used[k] |= xwith a defaultdict instead. - When writing generators that emit build inputs (stemlib files), make re-runs idempotent by merging with the existing output; otherwise entries generated in run N are dropped in run N+1 once the tool's own output becomes "recognized" input. Use an explicit managed-section marker so curated content is never rewritten.
- Managed-section trap (hit for real, Round 2): a generator that rewrites its
whole output file from its own lists will silently wipe anything hand-added above
the marker. Keep all curated entries in the script's data structures — never
hand-edit the managed region of the generated file. When normalizing emitted
values (e.g. grave→acute), apply the normalization at both entry points (new
data and
load_existing()merge) or re-runs won't converge.