MBIRA’s Evolve mode is a deterministic, corpus-constrained research model. It does not train on recordings, invent pitches, estimate a traditional grammar, or simulate a Zimbabwean musician. It starts with the selected source exactly and moves through complete physical-key cells borrowed from one project-reviewed sibling variation. Here, project-reviewed means manually registered and technically checked against the pinned source grids; it does not mean review by a Zimbabwean performer or lineage authority. Generated frames are temporary; stopping playback, changing tune, or disabling Evolve returns to the canonical source.
The player opens on Baya Wabaya EA Kushaura 1, a verified endpoint whose first changed phrase begins at the start of the second eight-second cycle, with Evolve selected. Choosing any of the 13 eligible endpoints also selects Evolve and Loop by default; choosing any other source falls back to a fixed cycle. The listener can still turn Evolve off. This policy makes the constrained model the normal experience where its exact structural vocabulary exists without pretending that every source has a safe variation path.
The narrow design follows from the evidence. The 34 source grids in mbiramachine are static loops rather than performances. They show related source states but not the probability, timing, or meaning of transitions between them. Seven registered sibling relationships provide 13 eligible endpoints and require the same recorded role state. Five endpoints across the Nyamaropa and Mahororo relationships retain the role `unknown` because their titles contain no unambiguous kushaura or kutsinhira token. The other source records stay fixed because a shared tune-family label is not enough to admit a transition.
The source hull is an explicit registry
The frozen version 1 recipe names seven bidirectional edges: Bukatiende Kushaura 2a and 2c; Nyamaropa Basic and Variation; Marenje Kushaura and Kushaura 6; Mahororo Basic and each of two variations; Chipindura Kushaura 1 and 2; and Baya Wabaya Kushaura 1 and 2. Each edge records which hand may change. Kushaura and kutsinhira are never made interchangeable by a title search.
The recipe also pins the context needed to interpret an endpoint: composition ID, exact tune IDs, role state derived from the source title, cycle size, pulses per beat, four phrase spans, hosho cells, part IDs, event duration, instrument type, and the complete physical-key-to-MIDI projection. The role is `kushaura` or `kutsinhira` when exactly one corresponding title token is present and `unknown` otherwise. This is versioned evidence, not metadata inferred afresh whenever playback starts.
Active phrase cells produce a reversible path
Eligible catalog records divide into four equal phrase compartments because that is the structure encoded by this source catalog. Those quarters are not performer-annotated phrase boundaries or a general claim about mbira form. The evolving state is a four-bit mask. A zero bit selects the source endpoint for that phrase; a one bit selects the donor endpoint. Bits whose two endpoints are identical are inactive, so the engine never spends a generation “changing” a phrase that sounds and plays the same.
A complete path starts at `0000`, toggles one active phrase at a time, reaches the available donor mask, and then releases by traversing the chosen order in reverse. A lane with `n` active phrases therefore has exactly `2n + 1` structural states: three, five, seven, or nine. There are no full-donor padding states, and the identical source state at the boundary between ranked lanes is coalesced. The opening source state lasts one playback cycle, so the first evolved frame starts on the next cycle. Every later state is held for two cycles. Executable tests bound every eligible run of musically identical cycles to those two intentional holds.
- Render the selected source exactly for cycle zero.
- Enumerate every order of the active one-bit phrase changes rather than shuffling phrase positions.
- Build each intermediate cycle from whole source or donor cells at the same aligned pulse and hand.
- Reject any step that fails the versioned transition evaluator.
- Rank complete accepted paths, use the seed hash only for a full-score tie, fall back to phrase order on a rare hash collision, hold each state, then reverse the path back to source.
Hard gates decide legality
The transition evaluator parses every foreign input through the strict tune schema and fails closed. It canonicalizes part, event, and chord order while preserving cyclic phase. Event IDs are ignored because generated frames need fresh IDs; musical attacks, physical keys, durations, pitch projections, velocity, pulse, and part ownership are not.
| Gate | Why it exists | Failure it blocks |
|---|---|---|
| Pinned endpoint lattice | Bind an ID to the exact compiled musical grid | A familiar ID attached to forged events |
| Project-reviewed edge and direction | Limit movement to a manually registered sibling relationship | Cross-family or role-changing interpolation |
| Whole aligned phrase allele | Keep each phrase coherently on one endpoint | Pulse-by-pulse mixtures absent from both sources |
| Physical-key provenance | Trace every attack to a registered source key | Invented notes disguised by a valid MIDI pitch |
| Cycle, hosho, part, and tuning context | Require the endpoints to mean the same structural axes | Combining incompatible grids or playback maps |
| One changed phrase and hand | Make each structural generation bounded and attributable | Unscored simultaneous rewrites |
An exhaustive audit across the project-reviewed directional lanes evaluated every phrase-allele background and made 11,776 aligned pulse-cell checks. It found zero cells outside the corresponding source or donor endpoint. Runtime tests separately accept all 628 distinct genuine grow and release toggles while rejecting forged endpoint grids, internally mixed phrases, malformed operations, raw two-hand donors, and context drift.
Soft metrics rank paths without pretending to authenticate them
Several paths can pass every hard gate. The evolution engine ranks complete paths lexicographically. Across proper intermediate growth states, it first minimizes the sum of phrase-boundary seam excess and then the maximum seam excess. Across all accepted non-no-op growth and release steps, it next minimizes adjacency novelty, onset salience, and total changed cells. Only when that complete musical score vector ties does a stable seed hash choose among paths; phrase order is the deterministic fallback for a rare hash collision.
Seam excess is calculated in the editorial MIDI register by comparing a candidate boundary with the worse boundary already present in its two endpoints. It can distinguish orders for Bukatiende, Mahororo Variation 2, Baya Wabaya, and Chipindura. It is deliberately not a gate. Chipindura’s best valid paths still have a nonzero total seam excess of 2.5, so rejecting every nonzero seam would discard legal endpoint-derived paths and turn a convenience metric into an authenticity claim.
Adjacency novelty and onset salience have the same status. They help prefer a less disruptive route among already legal whole-phrase states. They do not define Shona melody, hand technique, or ensemble fit. Grupe’s variation hierarchy supports gradual one-hand passage change as a conservative engineering direction; it does not supply transition probabilities for these catalog pairs.
Playback and visualization share one generated cycle
The model would still be misleading if the piano roll showed one state while the scheduler played another. The performance API materializes any absolute cycle directly from the source, selected lane, state hold, and deterministic path. The browser scheduler requests that same cycle for a short half-open time window. Lag recovery asks which generated cycle is currently audible before recovering sustained notes. The display switches only when the audible cycle index changes.
Before playback, the audio engine prewarms the union of every pitch reachable from the source and project-reviewed donor for the enabled hands and current transposition. Stop and parameter changes clear queued and active voices. Evolution implies looping because its state exists across cycles. It is selected by default for eligible records and remains structurally unavailable for records outside the exact registry.
Published analysis defines candidates; performance defines acceptance
Scherzinger’s counterpoint analysis and Azim’s teaching account explain why source-cell legality is not musical sufficiency. Scherzinger supplies analytical relationships that software can test; Azim shows that a player’s current focus determines whether two individually valid choices work together. Peterman’s interlock study adds explicit interlock categories and hand constraints. Together they support candidate analysis, not an automatic judgment of musical acceptance.
| Capability | Implemented output | What remains unclaimed |
|---|---|---|
| Inherent-line candidates | Several cyclic low-register-motion paths across multiple onset densities, with source event and part retained | Which line a musician or listener selects |
| Metric perspectives | Separate onset, duration, bass, harmonic-area, candidate-line, resultant, and hosho rotation profiles | One combined downbeat or meter winner |
| Interlock phases | Every cyclic offset with collision, alternation, coverage, high-register, and rest-complementarity descriptors | The performed or preferred phase |
| Variation vocabulary | Literal classification of accent, pitch, rest, octave, dyad, insertion, and exact metric-shift differences | Whether a classified operation belongs in generation |
| Harmonic and tuning evidence | Schemas for tune-relative areas, measured instruments, reference-only interval profiles, and explicit 12-tone error | Western chords or a tuning assigned by tune name |
| Performance sequences | A strict timed multi-cycle observation format with named performers, part-specific instruments, exact adjacent-frame differences, phase changes, expression, review, and annotation references | Transition probabilities or style claims before qualifying observations are supplied |
The candidate-line extractor never erases provenance when a line crosses hands or streams. Its output is ranked only by an inspectable cyclic registral-motion heuristic and spans several onset densities. The structural analyzer can report candidate-set continuity as a separate research diagnostic, but the evolution path does not use it as a gate because no line has been selected for the source performance.
Metric evidence is likewise kept plural. Hosho compatibility cells in the frozen recipe remain exactly what they were: endpoint metadata. A hosho perspective becomes available only when an annotation identifies landmarks for the exact material being analyzed. Harmonic-area evidence follows the same rule, using physical keys and tune-relative degrees instead of guessing Western chords.
The variation classifier describes literal differences before generation. It can recognize an exact cyclic shift or distinguish velocity accent, pitch insertion, rest substitution, octave movement, and dyad changes. It refuses to infer ghosting from low velocity and refuses to call a combined pitch-metric transformation conventional. A separate example annotation names its exact source and target spans, descriptive class, proposed generative operator if any, review, and acceptance question. That proposal is evidence metadata, never an executable grant.
The performance-observation boundary then asks a narrower question of adjacent cycles: what changed structurally, what changed only in articulation or intensity, did phase move, and do the declared part and pulse spans exactly cover the observed difference? Each part names its instrument, so paired performers are not forced onto one tuning map. Each transition becomes evidence-ready only with its own qualified independent review and references any selected-line or harmonic-area annotation it relies on. A review elsewhere in the session cannot make that transition ready.
Finally, the generative capability compiler separates evidence from software availability. `repeat-cycle` is always executable. `reviewed-phrase-allele` becomes executable only when its edge, direction, part, and span reproduce one exact entry in the frozen recipe and the selected source and donor cells materially differ inside that span. The remaining proposed operators can be described, but a reviewed annotation becomes evidence-ready only when its complete qualification record exactly matches a project-owned build registry; that registry is empty until such a record is retained. Even registered evidence stays unavailable until an implementation exists. A caller-authored edge ID, no-op phrase, annotation, operator label, or favorable review flag cannot widen playback.
Berliner and Magaya’s Mbira’s Restless Dance is the strongest identified next calibration source: 39 compositions and variations, 597 musical examples, Magaya’s performance advice, and companion audiovisuals. A future calibration stage should encode exact multi-cycle examples with their performer and instrument context, then compare transition density, phrase persistence, aggregate texture, selected lines, and phase relationships. A named, experienced Zimbabwean mbira musician must review any specific claim that the resulting behavior is player-like, traditional, or authentic.
A generative system becomes more credible when it can say no. MBIRA now says more precise kinds of no: a candidate is not a selected line, a descriptor is not a preferred phase, an observed difference is not an accepted operator, and a reference scale is not a measured instrument. Evolve remains a reversible walk through a small registered source hull; the analytical layer explains possibilities around that walk without silently widening it.