Skip to content

Latest commit

 

History

History
131 lines (120 loc) · 8.68 KB

File metadata and controls

131 lines (120 loc) · 8.68 KB

Lessons

Rules for future sessions in this repo, captured after corrections and mistakes.

Working files

  • All temp/working files go inside the morpheus folder (tmp/, git-ignored) — never in ~/.lmstudio/scratchpads. (User correction, 2026-10-07.)

Shell / tooling

  • The stemlib noun sources are named nom01…nom36 (no dot): the glob stemsrc/*.nom* silently matches only lsj.nom, nom.biblical, … Use stemsrc/nom* when searching all of them.
  • zsh chokes on unquoted parentheses in command lists; quote patterns or use a script file for anything with ( ).
  • When scrubbing a name from the repo, grep case-insensitively with word boundaries over ALL file types — not just .py/.md/.sh. Build artifacts (e.g. stemlib/Greek/conjfile, generated by concatenating stemsrc), generated reports, and saved snapshots all retain old strings long after the sources are clean. Fix at the root: edit the source, regenerate the artifact, delete stale snapshots.

Beta code / Morpheus internals

  • cruncher only matches lowercase beta code; uppercase input (reference- corpus style) silently returns nothing. The toolkit now normalizes this (runner._normalize_beta); keep it that way and test both cases.
  • to_beta() drops a breathing when the Unicode letter has none: plain Ι gives *israhl, but Ἰ (U+1F08) gives *)israh/l. When verifying name entries, test with properly breathed/accented forms AND check -n (accent-insensitive) mode.
  • Explicit :wd: noun/name forms match exactly (diacritics included); the accent-insensitive pass (cruncher -S -n) is what makes unaccented variants hit.
  • Generative :no:<stem> os_ou … patterns do not work for every stem shape (worked for σάββατον, produced nothing for τρυβλίον/ἀδελφιδός). Verify each new entry against the built index (steminds/nomind) and fall back to explicit :wd: lines.
  • The incremental stemlib makefile does not always pick up a brand-new stemsrc file; if entries are missing from nomind/vbind, check timestamps and force the target (e.g. touch stemsrc/lsj.nom && make).
  • Paradigm declarations (;xx, @ xx) in vbs sources never reach the built index. The build is cat … > conjfile; do_conj; mv conjfile.short /tmp/vbmorph; indexvbs — indexvbs indexes the do_conj output, and only processes :vs/:aj/:no/:vb/:wd/:de: lines (all ;, @, --prefixed lines are ignored). So ;pp/;pf declarations change nothing in vbind; use explicit :vb: lines.
  • The generative deriv path rejects non-reduplicated perfect stems (derivio.c checkcomderiv2: if (!had_redupl && Is_perfect(stemtype)) continue;), so -έω pp participles (εὐλογημένος) never generate. Explicit :vb: lines have no such gate.
  • GenIrregForm requires a real stemtype in the keys: o_stem alone is derivtypes-only and fails with "no stemtype seen"; add e.g. omi_pr before it (precedent: existing :vb:di/dw omi_pr …). FixRecAcc/FixPersAcc preserve any accent already present, so a :vb:/:wd: line storing the corpus's exact surface spelling round-trips exactly through CheckGenWords.
  • standword() rewrites GRAVE→ACUTE (every position) and strips accents after the first on every input before lookup; CheckGenWords compares accents exactly in normal mode (morphstrcmp, only |≡i equivalence) but stripaccs both sides under IGNORE_ACCENTS. Consequence: a stored word-final grave can never match an accented input — store the acute spelling instead (or both spellings, as with αλληλουία). The reference corpus is internally inconsistent (e.g. νοσσιὰν and νοσσιάν for one form; λημφθήτω unaccented in beta but accented in its Greek column) — index every attested standardized spelling.
  • Noun lookup (chckindecl) always strips all accents from the key; nomind keys are accent-stripped and merge all full-form variants sharing a stripped key onto one line (morphstrcmp: |≡i, plus "s1 ends ⇒ equal" quirk). The lookup is therefore accent-blind; only CheckGenWords' final compare is accent-sensitive. Multi-entry lines are fine — the parser iterates all entries.
  • C# port has placeholder stubs (chckindecl.cs returns 0; CheckGenWordsFunc uses plain string.Compare, no comptab/accent logic) — don't assume C/C# parity when reasoning about behavior from the .cs files.

Corpus tooling (LXX annotation pipeline)

  • Morpheus.analyze_beta() returns TokenResult wrappers, not raw analysis lists (result.analyses). The runner-level method returns bare lists; the API level does not.
  • Reference-corpus parse codes are position-dependent by lexical class: verb codes are 5 chars (tense voice mood person number), nominal codes only 3 (case number gender). Any field-by-field comparison must branch on class AND use per-class minimum lengths, or most nominals silently drop out of the comparison.
  • Morpheus's 8-char parsing code contains dashes (3AAI-S--), so shape-based column detection ("parse = the dash-free [A-Z0-9]+ field") misparses generated files — it reads the parse code as the lemma. Parse generated output positionally.
  • When training disambiguation models on a gold corpus and then evaluating agreement with that same corpus, say so in the report: the numbers measure domain consistency, not absolute accuracy.
  • Never hardcode the reference-corpus location (path or name) in any file. Scripts take --corpus-dir / env LXX_CORPUS_DIR; generated reports use a neutral label. The user greps the working tree — committed files, tmp/ scratch, and build artifacts must all stay clean.
  • The reference corpus is UPPERCASE beta code with its own dialect quirks (Round 6): breathings follow the letter (A)PE..., not APE...), / is a plain syllable separator inside words (KURI/OU = κυρίου — substring greps like RIOU miss it), and \, =, | are accent markers used inconsistently for the same word (αὐτοῦ appears as AU)TOU=, AU)TOU\S, AU)TOU/S). Never assume a spelling: key everything through canon() (to_beta∘from_beta + lowercase), and remember one word can occupy several canon keys that must all be covered by any table.
  • The eval's "top lemma disagreements" list is misleading (Round 6): it sorts by total form frequency and prints only the first disagreeing line, so a rare quirk on a frequent form looks like a mass error (ὅτι read as ×4,044; actually 7/4044). Always verify with true per-form mismatch counts before sizing a "prize pool".
  • Joint-frequency fallback: score unattested (lemma, POS) pairs at a floor of one occurrence — never zero, never full plain frequency (Round 6): the reference's type scheme and Morpheus's POS disagree for the same lemma (αὐτός A vs RD), so zero kills correct candidates; full plain frequency over-credits them (σός(N) 0.432 vs σύ(RP) 0.838 left a gap bigram context flips). An unattested pair can never be the gold reading of any token, so the floor only removes wrong answers.
  • The Viterbi tagger keeps one candidate per coarse POS (first on score ties) (Round 6): feature-level ambiguity within a POS — masc/neut homographs, duals — is resolved arbitrarily, and candidates with empty parse codes pass field comparison vacuously (no comparable pairs ⇒ True). Break same-POS ties explicitly (per-spelling reference aggregates; the corpus's accent notation encodes readings) instead of trusting Morpheus's internal candidate order.
  • Analysis.lemma/.variants() are Unicode, frequency-table keys normalized Unicode (Round 6): compare via normalize_lemma(), never raw literal equality — from_beta can return NFD and precomposed Greek literals won't match it.

Python

  • Augmented assignment on a function call is a SyntaxError: d.setdefault(k, set()) |= x — use d[k] |= x with a defaultdict instead.
  • When writing generators that emit build inputs (stemlib files), make re-runs idempotent by merging with the existing output; otherwise entries generated in run N are dropped in run N+1 once the tool's own output becomes "recognized" input. Use an explicit managed-section marker so curated content is never rewritten.
  • Managed-section trap (hit for real, Round 2): a generator that rewrites its whole output file from its own lists will silently wipe anything hand-added above the marker. Keep all curated entries in the script's data structures — never hand-edit the managed region of the generated file. When normalizing emitted values (e.g. grave→acute), apply the normalization at both entry points (new data and load_existing() merge) or re-runs won't converge.