Manual: Direction and the AI engine

The two controls that look like they do something magical and explain themselves the least. This is what they actually do — and, in Part 3, how to write a sound profile of your own.

Written because a user asked for exactly this. Everything below is generated from or checked against the shipping code.

It covers both builds. StemMaster runs on Windows 10 and 11, and on Apple Silicon (M1 or later) on macOS 15 or newer; Intel Macs are not supported. Everything on this page works the same on either — where a file path differs, both are given.

Two engines, and they read the Direction box completely differently. Almost everything on this page depends on which one is running, so it comes first.

Built-in analysis engineOptional AI engine
On by defaultYesNo
Needs anythingNoYour own API key
RunsOn your machineMeasurements go to your provider
Reads DirectionMatches a fixed word listReads your sentence

If you have not been into File → AI Engine Settings and entered a key, you are on the built-in engine, and Part 1 is the one that applies to you.

Part 1 — the Direction box

The box along the top, next to Add Stems. It is optional. Leave it empty and everything still works.

On the built-in engine it is word matching. Here is every word.

No language understanding, and it does not pretend otherwise. It looks for these terms and applies a bias for each one it finds. Everything else in your sentence is ignored.

You do not have to memorise the table below. With the built-in engine selected, the button beside the Direction box opens the same vocabulary inside the app, grouped, and clicking a word types it into the box — so you can pick one and keep writing around it. Cues that only lift a named stem are greyed out when your session has no such stem. The table here is the one that also tells you what each term does to the mix.

Write any of theseand it does
warm vintage analog tape−1.2 dB high shelf at 9 kHz; +1.0 dB low shelf at 120 Hz; saturation +0.10
bright crisp airy open shimmer sparkle+1.5 dB high shelf at 8.5 kHz; saturation −0.02
dark moody dusty lofi lo-fi−2.0 dB high shelf at 8 kHz; −0.6 dB peak at 4 kHz; saturation +0.05
bass low end low-end sub heavy thickBass +1.0 dB; +1.8 dB low shelf at 70 Hz
punch slam hit hard snappy knockDrums +0.7 dB; glue compressor attack ×1.5
wide spacious huge big room cinematicstereo width +0.10
vocal voice singer lyricsVocals +1.2 dB
guitar gtr riffGuitar +1.5 dB
keys piano rhodes organKeys +1.5 dB
synth padSynth +1.5 dB
solo leadGuitar +1.0 dB; Keys +1.0 dB; Synth +1.0 dB
aggressive gritty dirty distort raw hardsaturation +0.15
clean smooth soft gentle polishedsaturation −0.08; glue compressor ratio −0.2
spacey dreamy ambient lush atmospheric wet reverb hazy ethereal washedstereo width +0.06; reverb mix +0.08; reverb decay +0.5 s
dry tight in your face upfront direct no reverbforces every stem dry
pan left and right pan left to right auto-pan autopan auto pan swirl swirling ping-pong ping pong bounce between moving pan panning back and forthauto-pan, depth 0.4 at 0.22 Hz

Ordinary endings count, so punchy finds punch, bassy finds bass, warmer finds warm. Case does not matter, and it only matches whole words — subtle will not be read as sub.

Cues add up. “warm and punchy” applies both. Two cues pulling the same way stack, which is usually what you meant.

Contradictions do not cancel — the later stage wins. dry forces every stem dry, so “lush but dry” ends up dry rather than somewhere in between. If you want a middle, say one thing and adjust by hand afterwards.

When the words are not the problem

A cue moves a whole label: guitar lifts every guitar in the session. That is the wrong tool when one guitar is a rhythm part and another is the solo, because they carry the same label and measure much the same.

Worth knowing before you reach for it, because it is the same grouping seen from the other side: stems that share a type are balanced together, as one instrument. A kit tracked as kick, snare, overheads and a room mic is placed as a kit, and the balance between those mics is left exactly as you recorded it — the engine will not push a room mic up to meet the kick. Two mics on a bass rig, or a doubled rhythm guitar, work the same way. That is what you want almost always, and the setting below is how you overrule it for the one track that should stand out of its group.

Vocals are the exception, from 1.1.30. Six mics of one kit really are one kit. A lead with doubles and ad-libs behind it is not one voice recorded twelve times, and placing the pile pushes the lead down by however much you stacked around it — by how many parts there are rather than by anything you decided about the lead. So a group typed Vocals is placed by its loudest track: the lead lands where it would land on its own, and the parts behind it keep the balance you recorded underneath it. Nothing changes when a session has one vocal — the two rules decide identically there.

The Engine Notes name the track the group was placed by and say how far that moved it. That name is worth a glance. The whole group sits where one track puts it, so a loud ad-lib or a printed vocal effect left on the Vocals type would be levelling the record; retype it and re-run. If you preferred how 1.1.29 did it, put a file called vocal-levelling.txt containing the word sum in %LOCALAPPDATA%\\StemMaster on Windows, or ~/Library/Application Support/StemMaster on macOS, and the Engine Notes will say it is in force and what it cost you. Delete it to go back.

That is what the box under each stem's type is for. StemMaster picks a starting level from what a part usually is — a guitar sits below the vocal because a guitar is usually rhythm — and nothing in the file name or the audio distinguishes a solo from a backing part. So you say it:

Settingand it does
Primarybrings that stem forward — up to +2.5 dB on its usual level for that track type
Normalthe default: the usual level for the type
Secondarypushes it back by the same 2.5 dB

It is capped, so a Primary stem never lands more than about a decibel above where a lead vocal normally sits. It changes the level the engine mixes to, so it needs a full AI MIX & MASTER rather than the quick re-mix — and every stem it moved is named in the Engine Notes.

Watch the status bar

Whatever you type, the bar along the bottom tells you what was understood, a moment after you stop typing:

Direction heard: warm, punchy

or, if none of it landed:

Direction: no recognised words in that. The built-in
engine matches set terms — see Help → Direction Reference.

That second message is not a failure. The master runs perfectly well without a direction; it means those particular words did nothing, and you can either pick words from the table or reach for the AI engine.

It remembers what you asked for

Since 1.1.31 the box is a conversation rather than a fresh brief every time. Master, type vocals a bit louder, master again — and the AI engine is told what you asked before, what it reported doing, and that this newest message is a change to the mix you just heard. Say too much, back it off next and it knows what “it” is.

After the first master with a direction the results panel invites the follow-up:

Follow-up: say what to change and master again — it
remembers what you just asked.

That line appears once. The second message replaces it with the conversation so far, which is what the panel lists from then on — so it is there at the moment it is useful and gone as soon as you have acted on it. It does not appear on the built-in engine, which reads only your latest message and would not honour the invitation.

File → Clear Direction Conversation starts a fresh one and deliberately leaves the text in the box alone; someone who wants a new brief should not have to master once with an empty field to get one. The conversation is saved with the session and comes back when you reopen it.

Three things worth knowing before you rely on it:

What does not work, and why

make it sound expensive, like the last record, more vibe — the built-in engine finds nothing in any of these. There is no cleverness to appeal to. That is what the AI engine is for.

Part 2 — AI Engine Settings

File → AI Engine Settings, or Ctrl+,. Off by default. The built-in engine is a complete product on its own — this is an alternative decision-maker, not an unlock.

What changes when you turn it on

The decisions change, not the processing. Both engines drive the same EQ, compressors, reverb and limiter. The difference is what gets asked for.

What leaves your machine

This is the question behind the question, so: your audio never does.

What is sent is a JSON document of measurements. For the four-stem demo session it is about 3 KB. Per stem it carries duration, integrated LUFS, peak and RMS, crest factor, spectral centroid, band energy fractions, stereo width and correlation, dryness, band dynamics, and a count of samples pinned at the stem's own peak — plus the same for the summed mix, and the section map.

That last one is the only measurement here that is not a linear-domain quantity. It answers a question crest factor cannot: saturation, compression and clipping all crush the crest, but only a ceiling pins consecutive samples flat. It exists so the engine can tell deliberate coloration from damage — and leave the deliberate kind alone.

And one field you typed rather than measured: each stem's emphasisprimary, normal or secondary. It goes with the rest, because it is the one thing that tells the engine what the numbers cannot.

It also sends your stem file names, because they are often the strongest clue to what a track is. If your filenames carry a client or project name, that goes with them. Rename them first if that matters to you.

No samples, no waveform data, no audio in any form. That is not a promise, it is an observation — the outbound document was captured with the network call stubbed out:

size            2,940 characters
free text       stem name and label. Nothing else.
"audio"         absent
"waveform"      absent
"pcm"           absent
"sample"        present ONCE, as the key "samplerate"
longest list    4 items (the four stems)

That last line is the one that settles it. An audio buffer would be a list of hundreds of thousands of numbers. The longest list in the whole document has four entries in it.

The file that says what it decided

Separate from any of that, and true whichever engine you use: every master you export is accompanied by a text file with the same name, ending .settings.txt. It is written to your own disk, next to the audio. Nothing sends it anywhere; it is there so that you can.

It exists because of one support session. A customer heard low-end rumble in a master his own mix did not have, and could not have been more helpful about it — he sent his stems, his mix and our output. Diagnosing it still took a day and a half, because the file he sent had already been corrected by hand and nothing anywhere recorded what the app had decided before he corrected it. One line in one file would have answered it in a minute.

The same capture treatment as above, on the six-stem demo session:

size            7,161 characters
stems listed    6, each by filename and label
"waveform"      absent
"pcm"           absent
"audio"         present ONCE, in the sentence "It contains no audio"
longest list    6 items (the six stems)

It holds the profile and engine, the loudness target and sample rate, any direction you typed, the level, pan, EQ, high-pass and compression the engine chose for every stem, where your own faders ended up, and what the finished master measured. The top half is the summary panel's own words, copied rather than rewritten, so the file cannot quietly disagree with what you saw on screen. The bottom half is the same information as JSON between SETTINGS:BEGIN and SETTINGS:END, so the app can read it back.

Two things it does not contain: audio, and anything identifying you. It does carry your stem filenames, for the same reason the measurements document does — if those carry a client or project name, that is what you would be sending.

If a master ever sounds wrong to you, send this with it. It turns “something is off in the low end” into a question that can be answered.

It is your key and your bill

Which should you use

Use the built-in engine unless you have a reason not to. It is the default because it is good, it is instant, it costs nothing, and it works with no account anywhere.

Reach for the AI engine when the direction you want to give is a sentence rather than a word — arrangement-aware moves, per-section intent, or anything where “read what I actually wrote” is the point.

Whichever runs, the Engine Notes panel names it at the top and lists every decision it made. Nothing happens that you cannot read back.

The first line is Levels, and it is the only one that lists every track, including the ones left where they were — those read 0.0 dB. Every other line describes processing, so a track with no reverb needs no reverb line; this one answers “where did it put my bass”, and a missing entry could not tell you whether the answer was “nowhere” or “unchanged”.

It is also how you check whether a direction landed. Master, read the number, ask for the change, master again, compare. Added 1.1.33, after a session where “the bass at 1:03 is too overpowering” was answered by a render whose notes could not say whether the bass had been turned down — thirteen kinds of processing were enumerated and level was not among them.

When several tracks share a type

A kit arriving as separate mics is all one type — Drums — and until 1.1.28 that is what every line said, five times over, with no way to tell which drum got which setting. Attack of +3 dB is an ordinary choice on a kick and a mistake on a hi-hat, and they read identically.

Tracks sharing a type that were treated differently are now named:

Punch: Drums (01_Kick) attack +3 dB / sustain -2 dB;
       Drums (07_Hat) attack +1 dB / sustain -2 dB

Tracks given the same treatment still collapse to one line with a count — a stack of fifteen backing vocals on the same reverb is Vocals plate 8% ×15, not fifteen rows. One track per type, which is most sessions, reads exactly as it always did.

The sidechain key is the exception to “only when they differ”: it is named whenever your kit is separate mics, because which track fires the duck is the whole of that decision rather than a detail of it. From 1.1.28 the rule that refuses a tom, hat, snare or overhead for it applies on the AI engine as well as the built-in one, and when the AI's choice has to be corrected the line says what it replaced:

Sidechain: Drums (02_KickOut) → Bass 3 dB (re-keyed from 10_Tom2)

From 1.1.38 the duck is also gated on whether the two are actually in each other’s way. Two stems can both be low-heavy and never collide, because a bass often plays in the gaps between the kicks rather than under them — so StemMaster measures how much of the low end the bass holds at the kick’s own hits, and how much of it lives away from them. The second number is what separates a bass line from a distorted copy of the kick, which scores as a perfect collision and must never be ducked against its own source. The gate can only withhold: no session gains a duck it did not have before.

Part 3 — Writing your own profile

The PROFILE dropdown, next to Direction. Six built-in entries, and from 1.1.9, however many of your own you care to write.

Why this exists

Asked for by a tester on 10 August 2026, and the argument was unanswerable: feed StemMaster the full multitrack of a classic record and you get a clean master of a record that is not clean. That is not a bug. The engine aims at a modern commercial target — clear low end, controlled dynamics, translation across systems — and converges exactly where it is pointed. Point it at Motown 1971 and it will still take you to 2026.

The six built-in profiles are tonal destinations, not genres, and they were never going to become a map of recorded music. Knowing what a 1996 boom bap record is actually doing is not our knowledge to encode. So the format opens up.

You do not have to write it by hand

The easiest profile is the one you make with the faders. Master a track, move the faders and W sliders until it sounds right, then File → Custom Profiles → Save This Mix as a Profile. The difference between your mix and the engine's — your moves against the amber marks — is written out as a profile: your fader moves become stem_gain, your width sliders become stem_width, and the free-text box becomes the llm_prompt that steers the optional AI engine. The dialog shows in plain language exactly what will be saved before you save it, and a nudge too small to be a decision (under 0.3 dB, or one notch of width) is left out rather than enshrined.

The profile you were listening to is saved with it. Your moves were heard with that profile already applied, and the dropdown holds one profile at a time — so picking your capture replaces the one underneath rather than stacking on it. To stop that quietly changing the sound, the base profile's character (its master EQ, compression, width, saturation, punch and its AI-engine paragraph) is written into your file too, and its per-label levels are added to yours rather than replacing them. A capture made over Bright / Modern still sounds bright.

One limit worth knowing: band_targets is never captured, even from a base profile that has one. It steers the listen-back pass toward a tonal curve, and one song's arrangement is not evidence for that — a track with no cymbals would teach the profile to aim for no air on everything you master afterwards.

Where the files go

File → Custom Profiles → Open Profiles Folder. One .json file per profile. There is a worked example already in there — Motown '71 — to copy and edit rather than starting from a blank file. Save, then File → Custom Profiles → Reload Profiles, and it appears in the dropdown below the built-in six. No restart, so you can edit and listen without quitting.

Two halves, and you only need one

halfwhat it isneeds a key
the tilt master_eq, stem_gain, comp, width, saturation, punch, sidechain, space, stem_width, loudness, band_targets No
the steering llm_prompt Yes

The tilt is deterministic. It is applied after whichever engine decided the mix, so it works on the built-in engine with no key, no account and no network.

The steering is a paragraph appended to the AI engine's instructions, describing what that era or genre actually does. Only the AI engine ever sees it.

Either alone is a valid file. Describing a record in words and drawing an EQ curve are different skills, and demanding both would exclude most of the people worth hearing from.

Every field

{
  "name": "Motown '71",
  "description": "Midrange-forward, narrow, saturated.",
  "master_eq": [
    {"type": "highshelf", "freq": 9000, "gain_db": -2.5, "q": 0.7}
  ],
  "stem_gain": {"Vocals": 0.8, "Bass": 0.5},
  "comp": {"ratio_add": 0.3, "attack_mult": 1.0, "release_mult": 1.4},
  "width": -0.15,
  "saturation": 0.18,
  "punch": {"attack_add": -0.5, "sustain_add": 0.0, "parallel_add": 0.1},
  "sidechain": [
    {"key_label": "Drums", "target_labels": ["Bass"],
     "amount_db": 3.0, "release_ms": 180}
  ],
  "space": {"mix_add": -0.05, "decay_add": -0.4},
  "stem_width": {"Synth": 0.6, "Keys": 0.7},
  "loudness": {"max_clip_db": 1.0},
  "band_targets": {"air": 0.6, "presence": 0.9, "mid": 1.2, "sub": 0.7},
  "llm_prompt": "Cohesion, not clarity. Do not widen drums or bass…"
}

The same axis, by hand

stem_width is the written form of a control that is already on every channel strip: the W slider under the pan knob. Same axis, same rule — 100% leaves a stem alone, 0% folds it to mono, and it narrows only. On a mono stem there is nothing to take away, so it is greyed out with a tooltip rather than hidden, because a control that vanishes looks like a feature you do not have. Double-clicking it returns to 100%, not 0 — on a fader zero means unity, on a width slider it means mono, and those are not the same idea.

Where unity is, on a gain fader. The travel is −24 to +12 dB, so 0 dB sits a third of the way down and not in the middle. From 1.1.39 a scale beside each fader marks it: 0 brightest, the two ends of the travel next, and the steps between faintly. The scale answers “where am I heading”; the box under the fader answers “where am I” to a tenth of a dB, and accepts typed values. A narrow strip hides the scale along with the L/R captions and keeps the box.

The ticks are positioned from the slider’s own travel rather than from a fraction of its height. A fader’s handle moves through the height minus the handle, so the obvious arithmetic is out by up to 11 px at the ends — which is exactly where anyone would check a calibration mark, and a mark that is nearly right is worse than none.

The slider and the profile multiply, they do not override. The engine picks a width per stem, your profile scales it, and your slider scales that — so a profile asking for 0.5 on the Synth with the slider at 50% lands at 25%, not 50%. This is why the number beside the slider reads the slider’s own position and the applied figure lives in its tooltip. Showing the product there instead reads as a control stuck on a value it cannot leave: after a master it would say 39% with the handle at 100, and no position on the slider would make it say anything else.

Width is applied when the stems are summed rather than in the master chain, so moving it re-mixes rather than re-masters — it is one of the fast controls, and you hear it live while monitoring A, not only in the render.

It cannot break a render

Every number in your file is clamped to the same range the engine's own decisions are held to. A slipped decimal cannot produce something the app would not have produced itself — "gain_db": 400 becomes 9, not 400.

A file that will not parse is skipped, named, and the reason shown. The app still opens, the other profiles still load, and the status bar tells you which file and why. One stray comma should not cost you the application.

The balance you ask for is the balance you get

stem_gain moves a label by the dB you write, and until 1.1.37 that was not quite true — not because the offset was ignored, but because of what happened after it.

The engine balances stems from their measured loudness, and then compresses each one. Compression takes level off, and how much depends on how peaky the stem is; a kick loses far more than a sustained sample does. Nothing gave it back, so the mix that came out was not the mix the engine had decided. Measured on a customer’s six-stem session:

stemlabelcompression moved it, before 1.1.37from 1.1.37
J - DRUM DISTDrums−3.99 LUFS+0.18 LUFS
J - HHDrums−3.49−0.05
J - KICKDrums−5.13+0.12
J - SAMPSample−1.47+0.02
J - SNAREDrums−6.10−0.04
J - VOXVocals−2.40+0.02
spread across stems4.63 dB0.23 dB

The match is made on RMS, where it is exact to two decimals; the column above is LUFS, which is RMS with a hearing curve on it, and the two agree to within 0.2 dB.

The snare was arriving 4.6 dB quieter relative to the sample than the engine had asked for. Since 1.1.37 every per-stem compressor gives back what it took, so it changes dynamics and not level, and the same is true of each band of the master multiband.

Which way that moves a given record depends on which stems are the peaky ones, so it is worth being plain that this is not a bass control. On the session above the kick and the snare were losing the most, and the master gained 21% more energy below 150 Hz. On StemMaster’s own demo stems it goes the other way: there the sustained sub is the least affected stem at −0.56 LUFS, against −4.50 for the shaker and −3.30 for the drum loop, so level-matching brings the percussion back up and the low end’s share of the master falls about 6%. Nothing added bass in the first case and nothing removed it in the second. Both are the same thing — the balance the engine decided, arriving intact.

The one thing that follows from it: the master arrives at the loudness stage with its transients intact, so the limiter has a little more to do than it used to at the same target. That is the correct amount of work for it to be doing.

What a profile cannot reach

A profile biases the mix and master. It does not change what the processing chain is — both engines drive the same EQ, compressors, reverb and limiter, and a profile cannot add a device that is not there. If your period sound needs something the chain does not do, a profile will get you closer but not all the way, and that is worth knowing before you spend an evening on one.

Haven’t tried it yet? The free demo is the whole application with exports switched off, so you can run your own stems through it and read the Engine Notes before deciding anything. Windows 10 and 11.