Activated Cloud
← App Store

HyperFrames Audio

Activated Cloud✓ Officialactivated/hyperframes-audio

No ratings yet0 installsv1.0.0Updated Oct 6, 2026● Unknown

Free · Apache-2.0

About

Mixes audio already placed in a HyperFrames video: fades and crossfades, track gain, volume and effect automation, ducking, the voiceover carve that makes a music bed give way to a voice, effect chains (EQ, compressor, limiter, gate, saturation, delay, reverb, chorus, phaser, bitcrush) and submix buses. Diagnoses bad audio by measurement, since you cannot listen. Use when a placed track needs mixing or fixing. Not for finding or making music, sound or voice (use media-use) or for clip timing and track layout (use hyperframes-core).

Media

Documentation

From SKILL.md · v1.0.0 · what the agent reads when it loads this skill9 files: SKILL.md, references/CREDITS.md, references/attributes.md, references/carve-and-buses.md, references/diagnosis.md, references/fx-registry.md…

HyperFrames Audio

A mix is a set of relationships, not a stack of processors. Two tracks that each sound right alone can be unlistenable together, and the fix is almost never "turn one down": it is finding what they are fighting over and giving it to whichever one needs it. You finish a mix the way a good audio post engineer does: diagnose by measurement, subtract before you add, carve the bed under every voice, set a ceiling last, and prove the result with numbers from the rendered file.

Effects live on the element as data-fx-chain, and preview and render run the same Web Audio graph (a live context in preview, an offline one inside the browser the renderer already drives). There is one implementation of each effect, so what plays while scrubbing is what gets written.

When to use

  • "The music drowns out the voiceover." / "I can't make out the words."
  • "Fade the music in and out", "crossfade these two tracks", "duck the music under the narration".
  • "The voice sounds boomy / muffled / harsh / amateur", "there's a hum", "you can hear the room between sentences".
  • "Make it sound like a phone call / old radio / a big hall."
  • "Put one compressor on all the narration clips."
  • Before rendering any HyperFrames video that has music under a voice: the carve is part of finishing the mix.

What you need

  • The composition folder (index.html plus its audio and video files), on your computer in ~/Desktop/<your name> - Work space/<job>/.
  • Node 22+, FFmpeg and HyperFrames on your computer. Run the check in references/setup.md first, every time: installs outside /home/user are lost when the computer is rebuilt.
  • For the carve: @hyperframes/core installed in the project (npm i -D @hyperframes/core) and scripts/carve.mjs copied into the job (see Method step 5).
  • The owner's intent for the mix: is the music the point (a montage, a music video) or the support (a bed under narration)? If it is not obvious from the brief, ask with clarify.
  • The composition contract (data attributes, clip timing): load hyperframes-core with skill_view if you are unsure how clips and tracks are laid out. Sourcing new audio is media-use.

Three attributes carry everything, on the <audio> / <video> element itself (the first two also work on an <hf-audio-group> bus):

Attribute Holds
data-fx-chain the effects, in signal order
data-automation envelopes on this track's volume or its effect parameters
data-fx-carve the carve's own settings, so it can be re-derived

Exact JSON for each and the rules a lane must satisfy: references/attributes.md. Every effect with its parameters, ranges and units: references/fx-registry.md. Presets, named jobs, one-knob profiles and the full symptom table: references/presets.md (read it before hand-building a chain; a preset or job usually already names the problem). How to find out what is wrong with a file nobody has described: references/diagnosis.md. The carve, groups and buses in full: references/carve-and-buses.md.

Clip timing stays with hyperframes-core: trims and source ranges use data-start, data-duration and data-media-start, and a crossfade overlaps clips on different tracks. Constant data-playback-rate (0.1 to 10) is render-safe for picture and pitch-preserved sound when the matching audio and video elements use the same timing, source offset and rate. A speed ramp is a rate lane in data-automation; it wins over the constant and keeps pitch in preview and render. HyperFrames does not provide automatic waveform sync or drift correction.

Method

  1. Work out what is wrong before touching anything. You cannot listen, so measure. The governing rule: the absolute spectrum of a single unknown voice cannot be diagnosed. Formants are about ±10 dB, fundamentals run 85 to 255 Hz, and sentences decline 5 to 6 dB as they end; each of those reads as a defect on its own and each is the speaker. So compare against something inside the same file: the clean original if it exists, otherwise the pauses (anything audible in a gap was added, and the gap's spectrum is the channel, not the voice). Never compare against a published average spectrum or a different voice. If there is no original and no usable silence, a static tonal defect is under-determined: say so and give the owner the two or three readings that fit. The FFmpeg recipes (proportional 1/3-octave bands, noise floor, level over time, pitch) are in references/diagnosis.md.

  2. Name the symptom, then reach for its shipped answer.

    It sounds like Reach for
    Hum or thump underneath rumble-cut, or a highpass at 80 Hz
    Boomy, chesty Tame Boominess job (200 Hz)
    Muffled, behind cardboard Reduce Mud job (250 Hz)
    Words hard to make out Add Clarity job (3 kHz), or carve the bed
    Harsh and tiring Soften Harshness job (3.2 kHz)
    Some words much louder than others Evenness on a compressor, or Even Out Levels
    Room tone between sentences room-gate
    Voice and music fighting Voiceover carve, not an EQ on either
    Dry, recorded nowhere room-tight or room-natural
    Just "amateur" voice-clean, which is four of the above in order

    What is deliberately not covered (de-essing, noise removal, tone match) and the honest fallback for each: references/presets.md.

  3. Build the chain in the right order. Subtract before you add, level after you filter, relationships after level, character and ceiling last. Each step changes what the next one hears: a compressor set before a high-pass spends its time chasing rumble. The chain is serial, so corrective filtering goes early, character in the middle and a limiter last where it can act as a ceiling. Reach for a family by the problem:

    • Filters (highpass, lowpass, peaking, lowshelf, highshelf) decide which frequencies a track may occupy. They are the first tool for two sources colliding, because collisions happen in bands: a bed and a voice both want 1 to 3 kHz, and taking that from the bed costs far less than turning the whole bed down.
    • Dynamics (gain, compressor, limiter, gate) decide how level behaves over time. A compressor narrows loud-to-quiet; a limiter is only a ceiling; a gate removes what is below a threshold (room tone between phrases); gain is the plain stage an automation lane rides.
    • Nonlinear (saturate, bitcrush) adds harmonics that were not there. It is generative: it makes a thin source denser, not cleaner.
    • Time (delay, reverb, chorus, phaser) puts a track in a space. These wreck mixes most easily because a tail occupies the room a voice needs. Use them on what sits behind something else, with the wet amount lower than sounds right in isolation.
    • Give every node a stable id and, when it does a named job, a label ("Reduce Mud"), so lanes can address it and a reader can tell two peaking nodes apart.
  4. Fades, crossfades and ducking. A fade is a volume lane on the clip, in clip-local seconds (t: 0 is the clip's own start). A lane holds its first value back to the clip start and its last value forward to the end, so a bed that begins before the voice needs an explicit full-level point at t: 0. A crossfade is two overlapping clips on different tracks with opposite volume lanes. A plain duck is a gain or volume envelope; under a voice, prefer the carve.

  5. Carve every bed that plays under a voice. It is required, not polish. Put the voices in one group (data-audio-group="voiceover"), keep the bed and any SFX out of that group, then copy the script in and run it:

    • skill_view with name hyperframes-audio and file_path scripts/carve.mjs, then write_file it to <job>/tools/carve.mjs.
    • export PATH="$HOME/.local/node/bin:$PATH"; cd <job>/composition && npm i -D @hyperframes/core && node ../tools/carve.mjs --comp index.html --dry-run, read the report (which track it took as the bed, which as voices, the bands), then run it again without --dry-run.
    • Default strength 0.8 is right for narration; lower it only when the music is the point and the voice is sparse. Never carve a voice track. Carve against the group, never a list of clip ids. Rules and reasons: references/carve-and-buses.md.
  6. Use a bus when the same treatment belongs on several tracks. <hf-audio-group id="voiceover" data-fx-chain=... data-volume=...> gives every member one chain, one fader and one automation clock, and a compressor on the bus hears the whole voice. Bus automation runs in composition time, not clip time. A carve never goes on a bus.

  7. Automate only what can move. A lane targets volume or fx.<nodeId>.<param>. Only parameters backed by a Web Audio AudioParam move; compressor, limiter, gate and bitcrush expose none, so a lane on them is silently inert. To change a compressor's effect over time, automate a gain stage before it. references/fx-registry.md marks every automatable parameter. Read node ids back from the chain rather than assuming them: a lane pointing at a missing id is pruned without an error.

  8. Check, render, measure. Run npx hyperframes lint and npx hyperframes check (the static checks catch a volume lane fighting a GSAP volume tween, an authored data-volume on a tweened track, an ungrouped carve and a carve on a bus; nothing statically validates the chain itself). Render on your computer (load hyperframes-cli with skill_view for render flags), then measure the rendered file with FFmpeg (see Checks). A chain the renderer cannot parse fails the whole mix rather than writing a dry track; preview plays it dry. Effects with a tail (reverb, delay) make the rendered track longer than its source, which is expected.

Output

  • The composition with the mix written into its attributes (and any <hf-audio-group> buses), passing lint and check.
  • A short mix note for the owner: what was wrong (with the measurement that showed it), what you changed per track, the carve strength, and the final loudness figures. Example:
    • "Narration: rumble below 80 Hz (pause noise floor -52 dB) cut with rumble-cut; voice-clean for evenness. Music bed: carved against voiceover at 0.8. Final mix: -14.2 LUFS integrated, -1.3 dBTP peak."
  • The rendered video handed over by path and with show_card (type media) so the owner can play it.

Checks before you finish

  • Integrated loudness and true peak of the rendered file: ffmpeg -hide_banner -i out.mp4 -vn -af ebur128=peak=true -f null - 2>&1 | grep -E '^ +(I|LRA|Peak):' (I is integrated loudness, Peak is true peak). Many platforms normalise playback to around -14 LUFS and a true peak at or under -1 dBTP is a safe ceiling; confirm the current figure for the destination platform with web_search and state the source.
  • Under speech the bed sits clearly below the voice (as a rule of thumb, around 15 to 20 dB lower while the voice speaks) and comes back up between phrases. Check with windowed volumedetect on the voice and bed stems over the same seconds.
  • Every music bed under a voice has a data-fx-carve pointed at a voice group, and check reports no audio_carve_ungrouped_sources or audio_group_carve_attr issue.
  • Every lane's fx.<id> exists in that element's chain, and no lane targets a worklet effect.
  • A limiter, if used, is the last node of its chain.
  • If the carve sounds notched rather than simply quieter under the voice, the strength is too high: that is the one carve failure with an obvious signature, so look for it in the band measurements.

Pitfalls

  • Diagnosing one voice against a published curve or another voice. Two speakers differ by more than most defects. Compare inside the file, or report that it is under-determined.
  • Treating a smooth pause spectrum as proof there is no EQ problem. A filter multiplies, and near-silence times anything is still near-silence. Pauses reveal what was added, never what was filtered.
  • Ducking the whole bed instead of carving. The music goes limp for the whole voiceover. Carve, at 0.8.
  • Carving a voice, or putting the bed or an SFX in the voice group. The next analysis carves the bed against itself or ducks it under a whoosh.
  • Stacking a job on a preset that already contains it (voice-clean plus Reduce Mud is -6 dB at 250 Hz where -3 was meant). Expand the preset first.
  • Automating a compressor threshold. It cannot move. Automate a gain before it.
  • Levelling a track that only has natural declination. A 4 to 6 dB spread across sentences is normal speech; flattening it sounds robotic. Injected unevenness looks like 12 dB or more.
  • Shipping an explanation instead of a fix. When nothing shipped covers the problem (sibilance, say), name the gap, then apply the honest fallback and state its cost.
  • Inventing a clever new measurement to escape an under-determined answer. A confident result from a novel method, when both reliable references were unavailable, is the signature of the failure. Report the ambiguity and let the owner listen.

Versions

v1.0.0currentOct 6, 2026

Listed from the source repository.

Reviews

No reviews yet. Be the first.

Write a review