portfolio-press · applied-AI lab
internal productionlive audience on a channel

A melody hummed into a phone at night, and a scheduled video at the end of it.

No instrument, no click, no metadata: just a voice memo. What comes out of the other end is a mastered track, a cover, a video cut to the measured beat, subtitles landing on the syllable, and a publication booked on a channel. The line has done that 108 times. The one part that did not get faster is the moment a person listens and says yes.

108 · 164songs · masters alternate takes and versions counted separately, about 10.5 hours of finished audio.
~5songs a day at the ceiling what a good day now produces, from hum to finished master with its video.
~1 monthof publication queue the channel's schedule is booked that far ahead, so finished work waits its turn.
60download slots a month the one thing still rationed, and this month's are spent. More on that below.

What runs, end to end

A hum carries no metadata at all, so everything downstream has to be derived from the audio itself. Six stops, and a person enters at exactly two of them: once to choose a lyric, once to approve what ships.

01 · in

Arithmetic on a hum

Tempo, key, where the low end sits, how wide the image is. Derived from the recording, because nothing else knows.

02 · words

Three models, one human

The lyric goes around a loop of models with different failure modes. One writes smooth and sands off the awkwardness that makes a line sound like a person. One finds images. One argues about structure. The human picks, and that part is not automated.

03 · out

Generation, in threes

A style block built from measured production characteristics, never from someone else's melody. Back come siblings of the same song, and only one of them is going anywhere.

04 · repair

Stems and reassembly

A generated master is not a finished one. It comes apart into stems and goes back together with levels corrected, low end reshaped, and parts played over the top that the generator never played.

05 · picture

An edit grid from the audio

Cover art, then a video from stock, filmed and generated shots. Cuts a whole number of beats apart, accents on measured kick onsets, subtitles landing on the syllable. The grid comes from real beat positions, not from nominal tempo.

06 · out

An upload rail

Narration checked against its own script, metadata and thumbnail attached, and the result scheduled on the channel rather than posted by hand.

One track in the catalogue speeds up by 2.4 BPM per minute across its length. A grid built from its nominal tempo would drift off the beat by the last chorus, and nobody would be able to say why the video felt wrong. A grid built from the beats does not drift, and needed no opinion to get right.

The limit moved

This is the result worth reporting, and it is not "we make more music".

The limit used to sit on production: the hours between a melody in a phone and a finished master with a cover, a measured arrangement and a scheduled upload. It does not sit there any more.

A good day now produces about five songs worth finishing. The channel's publication schedule is booked roughly a month ahead, so finished work queues rather than ships. This month's sixty download slots are gone. Those are the same signal arriving from three directions: the making has outrun everything downstream of it.

Which is why the interesting question stopped being how to produce more. On a day with several finished tracks the one thing that does not scale is the part where a person listens and says yes. Every instrument described further down exists to make that moment happen more often, with better information, and with less of the listener's attention spent on things a machine should have counted.

What the machine does

  • Derives what a hum is actually at: tempo, key, balance, width.
  • Measures every candidate before a paid slot is spent on it.
  • Builds the edit grid from real beat positions and cuts to it.
  • Checks generated narration word for word against its script.
  • Schedules the publication and carries its metadata.

What it never decides

  • The melody, which belongs to whoever wrote it.
  • Which lyric survives the loop of three models.
  • Whether a finished track is good.
  • Whether anything is published. A person approves that.

The one thing still rationed

Everything above is cheap to repeat. One step is not, and it is the step where the choice happens.

The generative platform bills for every download and caps our plan at sixty a month. Those slots do not go on releases: a song arrives as two siblings from one request, then a remaster, then a regeneration after a lyric change, and every version you want to hear properly, on real speakers rather than in a browser tab, costs one. So the ordinary situation is three versions of one song, one credit left, and a decision that has to be right before it is spent.

Taste is not wrong here. It is just not sufficient on its own when the difference between two takes is subtle, you have heard both forty times, and it is four in the morning. So the line measures the boring things on every candidate first: tempo fitted to actual beat times rather than trusted from a library estimate, whether that tempo holds or drifts, how hard the kick sits on the beat, how much of the mix lives below 60 Hz, how wide the stereo image is, and whether any section is quieter than the rest. The measurements do not choose. They decide what gets listened to first.

One measurement turned out to matter more than the rest

Some generative models accept your own voice as a source. The question everybody asks is whether any of it carries through, or whether the model just notes the melody and sings in its own voice. The artist thought he could hear himself in one take and not in its sibling. That is exactly the kind of claim that cannot be settled by listening harder.

other artists' vocals, against the same reference 0.73–0.74 ordinary generation, no voice source supplied 0.78–0.82 the artist's own recordings, compared with each other 0.81–0.90 two takes from one request 0.784 the sibling 0.842 picked by ear one generation, then two remasters 0.808 first 0.759 0.735 after two the width of the method's own noise, 0.016–0.020, from two pairs 0.70 0.75 0.80 0.85 0.90
Speaker similarity, calibrated on our own material rather than on a published threshold. A similarity model runs locally on separated vocal stems. The three grey bands are the reference ranges that give the scale its meaning: other artists 0.73 to 0.74, ordinary generated vocals with no voice source 0.78 to 0.82, the artist's own different recordings measured against each other 0.81 to 0.90. The take he had picked out by ear scores 0.842, and the sibling take from the same session, which sounded fine to him but not like him, scores 0.784. He was right, and he was right before there was an instrument to check it with. The bracket is the measurement's own noise floor, drawn to length so that any drop on this scale can be compared against it by eye. It rests on two pairs of sibling takes, which is thin, and a floor established on two pairs is itself only roughly known. Reading across the rows is worth doing carefully: the twice-remastered take lands in the same region of the scale as other people's voices, which is an observation about where two published ranges sit, not a second finding.

The voice carries only where the singing went

  • Windowed across one track, similarity holds at 0.81 through the first two thirds.
  • It falls to 0.738 exactly where the melody climbs above the top note of the reference and half the pitches land outside the recorded range.
  • Practical form: if you want your own timbre on the chorus, sing the chorus, not just the verse.

Remastering erases it

  • One song, three versions of the same material: 0.808 for the first generation, 0.759 after a remaster, 0.735 after another.
  • The total fall is between 3.7 and 4.6 times the noise floor, so the direction of travel is real.
  • The second step is 1.2 to 1.5 times that floor, which is comparable with the noise rather than clear of it. It is reported here as one step, not as a trend that continues.

Why any of those numbers can be trusted

Because they were all wrong first. Every measuring tool written for this line produced confident numbers on real tracks, and nothing in the output looked off, because a wrong number and a right one are the same shape.

recorded · one evening, three defects

The control was a synthetic metronome, where the answer was known in advance

A tempo-drift tool was run against a generated click track whose tempo we had set ourselves. Three separate defects fell out of it, none of which had been visible on real material:

  • The underlying library reports tempo from a discrete grid of candidates and was off by two percent, which was enough to collapse a phase-concentration measurement to nearly zero on any material at all.
  • Estimating tempo independently in each window picked up half-tempo and triple-tempo readings, which read as drift of 22 to 35 BPM where there was none. The obvious fix, folding to the nearest octave, worked for the 2:1 case and broke on 3:2.
  • A separate onset tool reported times early by a constant 8.7 ms, and always missed the first onset of a file, because the peak picker needs history before it can pick anything.

The same discipline caught a written claim, not just a tool. The conclusion about remastering was first stated as a clean decline across three points, which is persuasive. Measuring the noise floor turned the second step from a trend into one step that is merely comparable with the noise, and the published version above says so.

the rule it produced Every measurement tool stamps its method into the result it emits, so that a measured value and an extrapolated one stop looking alike when they sit in the same sentence. Raw times are handed on raw, with a known offset as its own field, so two systematic corrections can never quietly merge into one that nobody can find again a month later.

The same control runs in the narration line, where it is cheap enough to leave on permanently: generated speech is transcribed back by a separate model and compared word for word against the script it came from. A divergence raises an alert instead of waiting to be noticed by ear. The daily forecast line uses the same gate, where one inserted word rejects the take.

What this page does not claim. The audience on the channel is real but small, and none of these numbers are about reach. Style prompts are written from measured production characteristics of a reference, tempo and balance and arrangement events, never from its melody. The similarity scale is calibrated on our own material, so it reads as a comparison inside this line and not as an absolute score anyone else should adopt. The quota and per-download price are the terms of our own plan in September 2026, not a claim about what any platform charges in general.

Why this sits on an engineering site

Because the habit transfers and the audio is incidental. A change that reads correctly in a diff is the exact analogue of a measurement that looks like a measurement. So the check is not the diff: the browser opens the running page and looks at it, or a script drives the interface and reports what happened, or the deployed version and the local one are compared directly and the difference is the review. It is the cheapest discipline available and it is almost never the default, because looking at the thing takes a minute and believing the diff takes none.

That is the same rule the rest of this site is built on, arrived at from a different direction: an engine computes, a model narrates, and a deterministic check re-verifies the narration against the engine before a human is asked for anything. The loop is drawn here, with the numbers behind it and the command that recounts them, and the gate can be watched refusing something without anyone's permission. The long version of this page's argument is an essay, Measuring What You Already Hear, with a companion inventory of what the line does end to end.

← the catalogue of productions you are here · from a hum to a channel next: the loop, drawn →