Bridging the Symbolic Gap: Moving MIDI to LLMs Without the Bloat
In REAPER + AI: Notes on a Data Problem, I looked at why feeding musical ideas to an LLM usually falls apart.
If you want an LLM to act as a theoretical sounding board—checking harmonic syntax or voice-leading intervals against a deadline—you run into an immediate systems question: How do you most effectively share musical time into text without destroying context or burning through token budgets?
The typical paths have obvious failure modes:
- Screenshots / Vision: Intuitively appealing, but vision models blur micro-timings, confuse vertical density, and hallucinate pitches. It adds an unnecessary extra layer of visual interpretation over data that is already digital.
- MusicXML: Standardized, but structurally massive. Layout data and engraving overhead consume thousands of tokens before the model ever reads a note. It is simply too much irrelevant information.
Both introduce latency and context starvation.
To solve this, I built a native in-memory bridge in REAPER (DN_Gemini_MIDI_Analyze_Selected) that skips disk read/write cycles entirely, focusing on three architectural choices.
1. In-Memory Direct Serialization
Instead of exporting files to disk, the script reads the active take's note buffer directly via REAPER's API. It strips layout, engraving, and container markup entirely, delivering a dense event array:
A multi-voice sketch that requires 4,000+ tokens in MusicXML drops to roughly 150 tokens. The context window stays focused on counterpoint, not markup parsing.
2. Managed Reasoning Budgets
With modern reasoning models, internal deductions and the final response share the same total output ceiling. Without explicit bounds, an LLM can easily burn its allowance on tangential internal scratchpad reasoning and truncate mid-cadence—or conversely, respond too shallowly.
By setting purposeful, balanced budgets on both sides (capping the thinking budget while providing a generous total ceiling), the model stays focused on the counterpoint without running out of runway. It makes every call predictable, affordable to run continuously, and consistently valuable.
3. Session Telemetry
To monitor the pipeline, the bridge logs local metrics (gemini_telemetry.jsonl): note density, dwell time, round-trip latency, and tokens-per-note. This turns a prompt experiment into measurable studio infrastructure, giving me long-term visibility to tune the pipeline as models evolve.
The Objective
Boundary control is essential: the model never generates or edits MIDI. It simply acts as an instantaneous second set of theoretical eyes inside the DAW.