honglab.

Decoding the neural architecture of behavior.

Science & Art

Zebrafish calcium imaging sonification: a five-stage project

You've just walked out of the imaging room with a hard drive full of freshly minted GCaMP6f traces from a Tg(elavl3:GCaMP6f) larva, your screen still glowing behind you, and you're already wondering…

Zebrafish calcium imaging sonification: a five-stage project

You've just walked out of the imaging room with a hard drive full of freshly minted GCaMP6f traces from a Tg(elavl3:GCaMP6f) larva, your screen still glowing behind you, and you're already wondering whether all those bursting voxels could become something a stranger would actually want to stand and listen to. Almost every lab I've mentored hits this exact wall around week three of a sonification pilot: the data is gorgeous, the publishable figure is still days away, and the leap from "ΔF/F₀ trace" to "audience-friendly sound" feels like a different PhD. It isn't. It's five stages, and once you've walked through them once, the next iteration gets quick. Let's go.

The hardest part of a calcium sonification project is not the imaging. It is making a clean, defensible decision about which number becomes which sound. Do that decision with intention, and the rest is craft.

Stage 1 — Optical acquisition: get the cleanest fluorescence signal you can

You cannot sonify garbage. Every artifact you tolerate at the microscope becomes an artifact you will eventually amplify through a speaker and explain to a museum visitor, so the upstream discipline pays for itself a hundred times over. For whole-brain larval zebrafish, the workhorse is a genetically encoded calcium indicator — GCaMP6f when you want fast, bright signals, GCaMP5G when you want a slightly better signal-to-baseline ratio and don't mind a touch less kinetics — expressed pan-neuronally under a promoter like elavl3 or HuC. Mount your larva in low-melt agarose on a raised platform, hold it in fish water with a gentle perfusion line, and keep the room dark. Bump up the infrared illumination only when you genuinely need to check focus; every photon you add is one your larva is paying for in phototoxicity, and a stressed fish is a noisy signal.

Light-sheet microscopy (SPIM) is still the right default for volumetric work, but extended light-field microscopy (XLFM) has been quietly reshaping what "whole-brain at speed" means in practice — recent implementations report volume rates up to 77 Hz across roughly one hundred thousand neurons simultaneously, which is more than enough bandwidth to give a sonification project room to breathe. If you're working with intact adult zebrafish where the skull scatters too much for two-photon, three-photon excitation (3PEF) opens the door to deep functional imaging in awake, behaving fish, because longer excitation wavelengths penetrate opaque tissue that would defeat a standard two-photon setup.

A practical note from the prep bench: pick your excitation pulse duration on purpose. The two-hundred-microsecond flashed excitation regime that some groups use is a deliberate tradeoff between signal yield and photobleaching, and if you're already planning to feed these traces into a sonification pipeline, the last thing you want is a slow, drifting baseline that you'll later mistake for musical phrasing. Lock down your settings now and document them — you'll thank yourself in stage 3. Whatever you do, note the behavioral state of the larva during acquisition (spontaneous, dark-reared, or lightly stimulated), because "what was the fish doing" is the first follow-up question every audience member will ask, and a one-word answer on the exhibit card is worth more than a paragraph in your notebook.

Stage 2 — Preprocessing and signal extraction: find the spikes in the noise

Let's get the boring-but-critical layer out of the way. Before anything touches a speaker, your raw light-sheet stack needs four passes: motion correction across the volumetric time series, source extraction (typically constrained non-negative matrix factorization, CNMF-E, or one of its faster cousins), ΔF/F₀ normalization against a rolling baseline, and finally optional spike inference if you want the sonifier to react to inferred spike events rather than raw fluorescence. You can absolutely skip spike inference and sonify the ΔF/F₀ trace directly — many of the most listenable pieces I've heard do exactly that, because raw fluorescence carries slow drifts that translate nicely into musical phrasing — but if you do skip it, you owe your audience an honest note in the exhibit label that the audio is a continuous proxy for calcium-bound events, not a spike-by-spike readout of action potentials. There is an important distinction: the underlying biology is fluorescence intensity from calcium-bound indicator, and the sonification faithfully reflects that, no more and no less.

This is also the stage where you decide what counts as a "neuron" for the purposes of your mapping. ROIs that survive your CNMF threshold become sonifiable units. Threshold too tightly and you will sonify only the loudest cells, which makes every soundscape sound like an anxiety attack. Threshold too loosely and you'll smear signal and noise into one indistinguishable texture and wonder why the output feels unfocused. There is no universally correct number; the unsatisfying rule of thumb is to set your thresholds so that your extracted ROIs look clean to a trained eye on a small sample of cells. If a postdoc in your lab can pick the true cells out of a random shuffled display in under thirty seconds, you're good.

Stage 3 — Parameter mapping: which number becomes which sound

This is the stage the whole project lives or dies on. A mapping is just a contract — you promise the audience that a specific measurable property of neural activity will always produce a specific kind of sound. Break that contract once and the exhibit loses its scientific story; keep it and you can stretch the artistic side almost anywhere you want. The conventional choices, and the ones I'd recommend as your starting menu, are these:

Neural propertyAudio mappingWhy it works
Peak ΔF/F₀ amplitude per ROIMIDI pitch or filter cutoffLoud cells naturally rise above quieter ones, like a melody over a drone
Burst onset timestampNote onset / triggerGives the listener a sense of event structure, easy to follow
Mean burst rate across an ROI populationTempo / rhythmic densityA swimming, behaving fish literally sounds busier than a quiet one
X/Y/Z position in the brain volumeStereo pan / binaural channel / ambisonic positionLets a viewer walk around and "hear" the brain the way they'd watch a map
Power in a chosen low-frequency band of the traceTimbre, harmonic content, or filter resonanceLets you talk about familiar brain rhythms in a way a layperson can actually feel
Number of co-active ROIs in a windowReverb send, ensemble loudness, densityA whole-brain flash event becomes a satisfying wash of sound

You do not have to use all six at once. Three is plenty: position, amplitude-to-pitch, and burst-to-onset is the trio I'd start with, because it gives a listener spatial, melodic, and rhythmic handles simultaneously. Pick the mapping, write it down in plain language, print it, and pin it next to the playback rig. If you cannot defend each row of that table to a curious ten-year-old in two sentences, simplify it.

Stage 4 — Auditory synthesis: making the mapped parameters actually sing

Once the contract is set, synthesis is mostly craft. Granular synthesis works beautifully for ΔF/F₀ traces because it can stretch a smooth fluorescence curve into a sustained, breathing texture; additive synthesis with partial tracking works beautifully for power-band mappings because you can literally voice the spectral shape; and if you're more comfortable in a tracker or a modular environment than in Max/MSP, Pure Data and SuperCollider both have well-trodden paths for exactly this kind of multi-channel mapping. MIDI mapping is your friend as a prototyping layer — a single neuron driving a single oscillator is two lines of code in any environment, and you can scale up to a thousand oscillators with relative confidence.

A few practical considerations from installs I've helped debug. First, mind your stereo field; if the listener is moving, binaural or ambisonic encoding is worth the extra setup, because a static stereo image loses its "I'm standing inside a brain" effect the moment the listener steps a meter to either side. Second, watch your gain staging at every junction: drop the level going into your final mixdown, and trust that your speakers do not need to be loud to feel intimate. Third, and this is the one I wish someone had shouted at me earlier: the larval zebrafish ear is tuned to roughly one hundred hertz to four kilohertz, so any frequency content you push well above 4 kHz is, sonically, an artistic choice with no biological correlate. You can absolutely make it, but write that choice into the label so the audience understands the artist is doing science, not pretending to be a fish.

If you want the audience to feel like they are inside the brain, give them spatial audio. If you want them to feel like they are watching the brain, keep the mapping strictly visual and let the sound ride gently on top.

Live, closed-loop sonification — playing back sound that reacts to optogenetic or behavioral triggers in real time — is the dream that everyone asks about and almost nobody gets all the way right, because the round-trip latency from stimulus to audio output is unforgiving. There is no widely cited real-time latency benchmark for this kind of loop, which is itself a useful warning: budget for iteration, and don't promise a curator a perfectly tight feedback rig until you've measured your own pipeline end to end.

Stage 5 — Public engagement: where the science actually leaves the building

Now the design problem changes. You are no longer optimizing for an N. You are optimizing for the thirty seconds a visitor has before their kid tugs their sleeve. Two pieces of advice that have saved every exhibition I've helped build. One: install a printed card, no smaller than the size of a hand, that explains in one sentence what the listener is hearing and in one sentence what biological signal maps to what sonic feature. Two: give the listener a choice between a one-minute and a five-minute loop. The one-minute loop is your hook; the five-minute loop is your reward for the person who stays.

Materials and budget for a small, portable rig are honest rather than exotic: a mid-range laptop, a multi-channel audio interface, a pair of good near-field monitors or a small ambisonic decoder rig, decent headphones for the listening post, printed labels, and a few dozen hours of your bench time spread across a month. That last line item is the one nobody budgets for. Sonification is rehearsal-heavy, and your first public mix will always be a little too busy. Three formats that have consistently worked in practice: a single-chair headphone booth with a printed card; a small floor-standing kiosk with directional speakers and a wall-mounted graphic of the same brain volume the audio is drawn from; and an open-plan ambient rig where the sonification plays into the gallery at conversational volume and visitors drift in and out. Each format has a different dwell-time budget, and choosing deliberately saves you from re-rigging three months later.

A short collaboration template that has actually worked

When you bring in a composer or sound artist, send them three things in advance: the explanatory card text, the mapping table from stage 3, and ten minutes of pre-rendered stems they can play with in their DAW before you meet. In return, ask them to produce two outputs — a finished stereo or binaural mix that conforms to your mapping table, and a "wild" version where they are explicitly invited to break the mapping in one well-marked section. The first version is the rigorous exhibit; the second version is the artistic response that lives next to it on the program. This pair tells the audience that the lab is rigorous and that the artist is trustworthy, which is the whole game at a public engagement event.

Closing: take the next prep this week

Pick one larva you already have good data for, sit with the ΔF/F₀ traces for fifteen minutes with the audio routing off, and just listen with your inner ear. Where do you hear rhythm? Where do you hear pitch? Where do you hear silence? That sketch is your real stage-3 mapping — the rest of the project is just making it honest. I still get a small nervous thrill every time a non-neuroscientist hears one of these pieces and says, "wait, go back, that bit near the end where it quiets down — what was the fish doing?" That is the conversation our field needs more of, and you are closer to it than your last imaging session made you think.

FAQ

What is the most important step in a zebrafish sonification project?
The most critical step is defining a clear, defensible mapping contract where specific measurable properties of neural activity are consistently assigned to specific sounds.
Can I sonify raw fluorescence traces directly?
Yes, you can sonify ΔF/F₀ traces directly, as raw fluorescence often contains slow drifts that translate well into musical phrasing.
How should I explain the sonification to a museum visitor?
You should provide a printed card that explains in one sentence what the listener is hearing and in one sentence which biological signal maps to which sonic feature.
Why is spatial audio recommended for these projects?
Spatial audio, such as binaural or ambisonic encoding, allows the listener to feel as though they are inside the brain rather than just observing it from the outside.
What should I do if my sonification includes frequencies above 4 kHz?
Since the larval zebrafish ear is tuned to a range of roughly 100 Hz to 4 kHz, any content above 4 kHz is an artistic choice that should be explicitly noted in the exhibit label.