BeatsToRapOn Field Guide
Karaoke Production · 2026
Complete workflow · audio separation · lyrics · performance

How to make a karaoke track from any song.

Turn a finished song into a singable backing track, then add synced lyrics, choose the right key and export it for rehearsal, a party, a video or the stage. This guide separates what works from what merely makes the original vocal quieter.

The 30-second answer

Upload the cleanest copy of your song to an AI vocal remover, download the instrumental, inspect the loudest chorus for vocal bleed, repair only the obvious artifacts, then add timed lyrics and export a lossless audio master plus the delivery format you need.

2→1

Start with vocals + instrumental. Use more stems only when the backing needs repair.

01 Define the deliverable

A karaoke track is more than “no vocals.”

A finished karaoke asset can contain up to four separate components: the backing audio, lyric timing, visual presentation and performance cues. Decide which ones you need before you start editing.

The instrumental is the foundation. It keeps the drums, bass, harmony and arrangement while reducing or removing the original lead vocal. A basic audio-only track may be enough for private practice. A karaoke video also needs readable, synchronised lyrics. A live performance file may need a count-in, key change, guide cues and a reliable ending.

No vocal remover can promise a perfect result from literally every mix. Lead vocals, backing voices, reverb, guitars, synths and cymbals often occupy the same frequencies and stereo space. Modern source separation estimates the hidden parts; it does not retrieve original studio stems that were never supplied.

LAYER / 01

Backing

The instrumental performance without the original lead vocal.

LAYER / 02

Lyrics

Accurate words split into readable lines and timed to the vocal entry.

LAYER / 03

Visuals

Contrast, highlighting and safe placement for a screen or video.

LAYER / 04

Cues

Count-ins, key changes, pickups and endings that support the singer.

Instrumental, karaoke version and backing track: what is the difference?

TermUsually meansMay includeBest use
InstrumentalThe music without the lead vocalBacking vocals, ad-libs or no guide cuesListening, remixing, practice
Karaoke trackAn instrumental prepared for a singerSynced lyrics, count-in, key change, guide melodySing-alongs, venues, video
Backing trackPre-recorded accompaniment for a live performerClick, cues, backing vocals or additional productionRehearsal and stage performance
Minus-oneA mix with one featured part removedEverything except voice, guitar, drums or another partMusic practice and auditions

Apple’s Music Sing is useful context: it offers adjustable vocals and real-time lyrics across millions of songs, but its control changes vocal level inside Apple Music rather than exporting a new backing-track file.[9] If you need an actual file for editing, rehearsal or video production, create and export the instrumental yourself.

02 Pick the right process

Six ways to make a karaoke version.

AI separation is the most practical starting point for a finished commercial mix. Other methods can outperform it when you possess better source material or need a legally clean re-recording.

MethodBest situationControlMain weakness
AI two-stem removerFastest karaoke instrumentalVocal + instrumentalSome bleed or backing loss can remain
AI multi-stem splitterBacking needs repair or rebalanceVocals, drums, bass and instrumentsMore files and mixing decisions
DAW stem separationYou already edit in a compatible DAWVaries by applicationSoftware and hardware requirements vary
Exact phase cancellationYou have the matching official instrumentalPotentially very preciseFails if masters, timing or gain differ
Centre-channel reductionOld editor and simple centred vocalLowAlso removes centred kick, bass and snare
Re-record the arrangementCommercial-quality custom backingMaximum musical controlTime, skill, cost and composition rights

1. AI vocal remover: the default choice

A two-stem model estimates the song as vocals plus accompaniment. It is ideal when the desired output is simply “the same song without the singer.” BTR’s AI Vocal Remover accepts common formats including MP3, WAV, FLAC, AAC and M4A, then provides vocal and instrumental outputs for preview and download.

2. AI stem splitter: the repairable choice

If vocal removal damages the bass, drums or harmony, separate the mix more deeply. The BTR AI Stem Splitter offers four- and six-stem workflows, including vocals, drums, bass, guitar, piano and other instruments. Recombine everything except the vocal, then rebalance or replace any damaged part.

3. Built-in DAW separation

Some DAWs now integrate source separation directly into a project. This can be convenient because the extracted regions land on editable tracks. The quality ceiling is still set by the mix, while availability depends on the software version and hardware. Browser processing is simpler when you only need an instrumental file.

4. Exact phase cancellation

When you own the released full mix and the exact instrumental used to create it, align them sample-accurately, match gain, invert the polarity of one and sum them. Shared information can cancel, leaving the difference—often the vocal. This is not the same as “removing the centre.” A one-sample offset, alternate limiter, different encode or revised master can prevent clean cancellation.

5. Centre-channel reduction

Older karaoke effects subtract information common to the left and right channels because many lead vocals are mixed near the centre. The method cannot distinguish a centred singer from a centred kick, snare, bass or lead instrument. Stereo doubles, reverb and delay often survive. Use it only when AI processing is unavailable or when the mix happens to suit it.

6. Re-record the music

A producer can rebuild the arrangement using new performances and instruments. This avoids source-separation artifacts and allows a custom key, length and arrangement. It does not remove the need to clear the underlying composition, lyrics or public use. It is the highest-control route, not the fastest.

Use the smallest tool that solves the job.

Start with two stems. Move to six stems only if the result needs surgery. Every extra output creates more control—and more opportunities to change the original balance.

03 Fastest complete workflow

How to make a karaoke track in eight steps.

This route creates an audio backing track first. Lyric video instructions follow in a separate section so the audio can be approved before hours are spent timing text.

01

Choose a legitimate, high-quality source

Start with the cleanest file you are authorised to use. WAV or FLAC preserves the source without additional lossy encoding. A high-quality MP3 or M4A can still work; a screen recording, speaker capture or repeatedly converted file gives the model less reliable information.

  • Use the full song, not a social-media excerpt.
  • Avoid normalising, clipping or limiting the file again.
  • Do not convert an MP3 to WAV expecting lost detail to return.
02

Separate vocals and instrumental

Open the BTR AI Vocal Remover, choose the audio file and start processing. The goal is two synchronised outputs: an isolated vocal and the accompaniment. Keep the browser tab open during upload and processing.

03

Preview the hardest sections

Do not approve the result after listening to the intro. Check the first vocal entry, the loudest chorus, backing-vocal stacks, a quiet verse, exposed breakdowns and the final reverb tail. Use headphones first, then ordinary speakers.

  • Listen for words: lead phrases, ad-libs, doubles and harmony residue.
  • Listen for missing music: snare attack, bass notes, guitars or synths pulled into the vocal stem.
  • Check stability: warbling ambience, pumping or a hollow stereo image.
04

Download the instrumental and vocal

Download the instrumental for the karaoke track. Save the vocal too, even if you do not plan to use it. The vocal output helps identify where missing instruments went, locate lyric entrances and compare timing later. Keep both files at their original full length.

05

Repair only the audible problems

Import the instrumental into a DAW or editor. Use volume automation to reduce isolated vocal fragments during exposed gaps. Add short crossfades around edits. If a small artifact is covered once the new singer performs, leave it alone; aggressive EQ and denoising can make the backing thinner than the artifact itself.

06

Choose the singer’s key before timing lyrics

Detect the original key and BPM with BTR’s Song Key & BPM Finder. Test the highest and lowest phrases with the actual singer. If the key changes, transpose the finished instrumental before synchronising lyrics so the timing and final audio remain locked.

07

Add a count-in, lyrics and performance cues

For audio-only rehearsal, a one- or two-bar count-in may be enough. For karaoke video, transcribe or license the lyrics, split them into singable lines, and time each line to appear before its first syllable. Mark instrumental sections and pickups clearly.

08

Export a master and test the whole performance

Keep a WAV master for future edits. Export a compressed copy only when the playback device, video platform or delivery channel requires it. Sing the entire song from the final file and on the final playback system. Confirm the first cue, key, lyric timing, level and ending.

  • Name clearly: SongTitle_KARAOKE_Key-WAV.wav.
  • Leave sensible headroom for a live singer and PA.
  • Carry a backup copy on a second device for live use.

Make the first instrumental now.

Upload your source, preview the vocal-free backing and download the result. No production software is required for the initial split.

04 When two stems are not enough

Build a cleaner backing from multiple stems.

A multi-stem split will not magically recreate the original multitrack session, but it gives you separate controls for the parts most likely to be damaged by vocal removal.

Split the source into vocals, drums, bass and instruments—or into a six-stem arrangement when guitar, piano and other parts need independent treatment. Import every output at the same start time. Mute the vocal stem, then compare the remaining sum with the two-stem instrumental at equal loudness.

STEM / 01

Drums

Protect impact, timing and cymbal detail; replace only damaged hits if necessary.

STEM / 02

Bass

Restore weight lost where vocal fundamentals and bass notes overlapped.

STEM / 03

Harmony

Balance guitar, piano and other parts without making the track feel hollow.

STEM / 04

Vocals

Mute the lead, or retain selected backing phrases only when appropriate.

A practical hybrid repair

The cleanest result may use parts of both jobs. Keep the two-stem instrumental as the base. Place the separate drum and bass stems underneath only where the base loses impact. Align every file from time zero, check polarity and avoid running them together at full level across the entire song; duplicated estimates can produce comb filtering or excessive low end.

  • Import all stems together. Never trim their fronts independently.
  • Level-match comparisons. Louder nearly always sounds “better” in a quick test.
  • Mute first, then rebuild. Add parts until the backing matches the musical energy of the original.
  • Automate by section. A chorus may need different repair than a sparse verse.
  • Check mono. Stereo tricks can hide phase problems until the file reaches a PA.
  • Render once. Keep a lossless session and avoid repeated lossy exports.

If you also want to reuse the extracted vocal creatively, the same files can feed a mashup or remix. Keep that project separate from the karaoke master so experimental processing cannot damage the performance version.

05 The source sets the ceiling

Why some songs make better karaoke tracks.

Separation quality depends on more than file extension. The arrangement, vocal effects, mastering and encoding determine how much evidence the model has for each hidden source.

Source featureUsually easierUsually harderWhat to do
EncodingOriginal WAV or FLACLow-bitrate or repeated MP3/AAC conversionReturn to the cleanest legitimate source
Lead vocalClear, stable, mostly centred voiceWide doubles, distortion and dense stacksTry multi-stem separation and automation
Vocal effectsShort controlled ambienceLong reverb, delays and chorusReduce exposed tails by section, not globally
ArrangementSpace around the vocalGuitars, strings or synths masking every phraseRebuild from more stems
MasteringClean dynamics and limited clippingHeavy limiting, clipping and saturationAvoid further limiting before separation
Backing vocalsClearly distinct from leadChoirs and harmonies spread across stereo fieldDecide whether to keep them; automate exposed words

WAV versus MP3: what actually changes?

WAV and FLAC can preserve the full decoded signal without perceptual data removal. MP3 and AAC discard information according to psychoacoustic models. A good encode may still separate well, but audible swirls, softened transients and high-frequency smearing can make source boundaries less stable. The practical rule is simple: use lossless when you genuinely have it; otherwise use the highest-quality original file available.

Converting a compressed file to WAV does not improve its source quality. It creates an uncompressed container around the already-decoded audio. The missing detail remains missing. Likewise, downloading or recording a streamed song can introduce another generation of conversion and may breach the service’s terms or the owner’s rights.

Why reverb is often the last “voice” left behind

Dry lead vocal may be centred and recognisable, while its reverb and delay spread across time, frequency and the stereo field. Those effects resemble part of the surrounding instrumental ambience, so a separator may place some of them in the accompaniment. Reduce an exposed tail with automation only when it distracts. In a room with a new singer, small remnants are often masked naturally.

06 Turn audio into karaoke

How to add synced lyrics and make a video.

Good karaoke lyrics tell the singer what is coming before they need to sing it. Accuracy, anticipation and contrast matter more than elaborate animation.

01

Prepare accurate lyric text

Use text you wrote, have licensed or are otherwise authorised to reproduce. Check repeated choruses instead of assuming they are identical. Preserve meaningful contractions, backing responses and language marks. Decide whether ad-libs are essential or distracting.

02

Split words into singable lines

A line should fit comfortably on the target screen and represent one musical phrase. Avoid placing half a phrase on a new screen merely to keep visual symmetry. Two short lines are usually easier to scan than one long sentence.

03

Mark the vocal entrances

Use the isolated vocal as a timing reference. Place markers at each phrase start, pickup and sustained final word. For line-by-line karaoke, bring the next line on screen roughly one musical beat before the singer enters. For word-level highlighting, align the highlight with the syllable, not the written word boundary.

04

Design for the worst screen

Use large type, strong contrast and a safe margin from every edge. Test on a phone, television and projected image if those destinations matter. Never communicate the current line using colour alone; position, highlight weight or a progress treatment should reinforce it.

05

Label instrumental gaps

Show “[instrumental]”, a countdown or a simple progress cue during long breaks. Mark duets, spoken parts and key changes. An eight-bar silence without context feels like a technical failure to a nervous performer.

06

Export and watch without singing

Review once as a singer and once as an operator. Confirm that every line appears early enough, stays visible long enough and disappears without covering the next phrase. Check audio-video synchronisation after the final encode, not only inside the editor.

Recommended karaoke video layout

Top / context

Next line

Show the upcoming lyric in a quieter state so the singer can prepare.

Centre / action

Current line

Use the highest contrast and a clear progress or highlight treatment.

Bottom / navigation

Cue bar

Reserve space for count-ins, instrumental breaks, duet names and key changes.

Rights reminder: lyrics are part of the musical work, not free interface text. Publishing them in a karaoke video can require permission even when the backing audio is newly created. See the copyright section below.

07 Fit the singer

Change the key and tempo without wrecking the track.

The “correct” karaoke key is the one the performer can sing consistently, not automatically the key of the original record.

Use the Song Key & BPM Finder to establish a starting point. Ask the singer to perform the highest chorus and the lowest verse over the actual backing. Move in semitone steps. A change of one or two semitones can be enough; a large shift may make drums, cymbals and formants sound unnatural.

ProblemAdjustmentCheck immediatelyRisk
Chorus is too highTranspose down 1–3 semitonesLowest verse notesLow instruments may become muddy
Verse is too lowTranspose up 1–3 semitonesHighest chorus noteCymbals and ambience may sharpen
Singer rushesReduce tempo slightlyNatural phrase endingsTime-stretch artifacts on transients
Song drags liveIncrease tempo slightlyBreath points and fast lyricsLess room for difficult phrases
Large key shift neededRebuild or re-record backingInstrument tone and vocal comfortSeparated artifacts become more obvious

Pitch shift first or separate first?

For most jobs, separate the clean original first, then transpose the instrumental once. Pitch-shifting before separation changes the spectral cues available to the model; shifting both before and after adds unnecessary processing. If a large transpose reveals artifacts, test both orders and use the better result rather than relying on a universal rule.

Keep the arrangement locked

When changing tempo, process the final backing as one file unless you deliberately need independent stem control. If separate stems are time-stretched with different settings, transients can drift and ambience can stop lining up. Render the chosen key and tempo before final lyric timing, then lock the audio.

08 Performance-ready delivery

Mix, master and export for the real room.

A backing track that sounds impressive in headphones can fail on a television, Bluetooth speaker or venue PA. Build for the destination and keep a clean master.

  • Leave headroom. Do not chase the loudness of the original master when a live microphone must sit over it.
  • Protect the low end. Check bass and kick in mono, especially after combining multiple estimates.
  • Keep a count-in optional. Make one stage version with it and one general karaoke version without it.
  • Use gentle fades. Preserve the song’s intended ending unless the performance needs a defined cut.
  • Test the first second. Some playback systems clip an immediate entrance or add Bluetooth latency.
  • Back up locally. Do not depend on venue Wi-Fi or one cloud account during a show.

Export settings by destination

DestinationMasterDelivery copyExtra check
DAW or future editingWAV, original sample rateNone requiredPreserve full-length timing
Live performanceWAVWAV or device-supported lossless fileTest on the actual playback rig
Karaoke videoWAV audio masterHigh-quality audio inside final videoCheck sync after encoding
Phone or casual partyWAV archiveHigh-quality MP3 or AACConfirm offline playback
Venue libraryWAV archiveFormat specified by the systemUse searchable artist/title/key metadata

Phone, browser or computer?

Phone · fastest start

Separate

Use the browser to upload and preview. Headphones help reveal vocal residue that phone speakers hide.

Computer · complete build

Edit

Use a DAW or video editor for automation, key changes, timed lyrics and final export.

Venue · final truth

Test

Play the exact delivery file through the actual system and sing the entire arrangement.

09 What the data really says

Source separation has benchmarks—not guarantees.

The research shows dramatic model progress, but it also explains why one headline score cannot tell you whether a particular chorus will sound clean.

150Full-length tracks in the widely used MUSDB18-HQ reference dataset
9.20 dBReported average SDR for Hybrid Transformer Demucs with extra training data
9.80 dBReported average SDR for a smaller BS-RoFormer on MUSDB18-HQ without extra songs

The dataset behind many separation claims

MUSDB18-HQ contains 150 full-length stereo songs: 100 in its training partition and 50 in its test partition. Each 44.1 kHz WAV song includes the mixture plus four source groups—drums, bass, vocals and “other.” Its maintainers describe source separation and karaoke among the dataset’s uses.[1] The collection is influential, but 50 hidden test songs cannot represent every production style, language, mix density or vocal effect a user may upload.

How model results moved

Open-Unmix and Spleeter helped establish accessible, reproducible systems for music source separation in 2019 and 2020.[2][3] Later hybrid systems combined waveform and spectrogram processing. A 2021 study reported a 1.4 dB average SDR improvement over its non-hybrid baseline on MUSDB-HQ and higher listener ratings for both overall quality and absence of contamination.[6]

Hybrid Transformer Demucs reported 9.20 dB average SDR on MUSDB when trained with 800 additional songs.[4] A smaller BS-RoFormer reported 9.80 dB average SDR on MUSDB18-HQ without extra training songs, while a larger system trained with 500 additional songs ranked first in the SDX23 music-separation track.[5]

Published resultReported dataWhat it supportsWhy not compare blindly
Hybrid model study+1.4 dB average SDR over non-hybrid baselineCombining domains improved objective resultsModel, dataset and baseline are study-specific
HT Demucs9.20 dB average SDRStrong four-source benchmark resultUsed 800 extra training songs
BS-RoFormer9.80 dB average SDRStrong MUSDB18-HQ result without extra songs for the smaller modelArchitecture, evaluation and source averaging differ
2025 listener studyAbout 30 ratings per track across seven listener groupsMetric usefulness differs by sourceNo single metric predicted every perceptual outcome

Important: these are results reported by separate research teams under their stated conditions. They are not BTR product accuracy claims, and they should not be treated as a direct leaderboard unless datasets, training data, source definitions and evaluation code match.

What does SDR mean?

Signal-to-distortion ratio is an objective measure used to compare an estimated source with a known reference source. Higher is generally better under the same evaluation setup. But the karaoke listener cares about specific failures: a recognisable word left in the chorus, a missing snare transient or an unnatural reverb tail. An average score can hide those local problems.

A 2025 preprint studying about 30 ratings per track across seven listener groups found that SDR best predicted listener ratings for vocals, while other measures were more informative for drums and bass. It found no tested embedding metric that correlated positively with human perception for vocal estimates.[7] An ISMIR 2025 study likewise argued that averaged separation metrics can obscure distribution differences and task-specific failure cases.[8]

The benchmark chooses a model. Your ears approve the song.

A defensible listening test

  • Compare outputs at matched loudness; do not reward the louder file.
  • Test at least one loud chorus, quiet verse, vocal stack and exposed instrumental gap.
  • Listen to the instrumental alone and underneath a new singer.
  • Check headphones, mono playback and the intended speaker system.
  • Write down the failure before processing: “word remains at 1:42” is actionable; “sounds AI” is not.
10 Before you publish or perform

Making the file does not clear the rights.

Vocal removal changes audio. It does not grant ownership of the song, recording or lyrics.

Private practice versus public use

Planned useQuestions to resolveLow-risk starting point
Private home practiceWas the source acquired and used lawfully?Use your own lawful copy and keep the result private
Public karaoke venueDoes the venue hold required performance and karaoke permissions?Use authorised karaoke catalogues and confirm venue coverage
Online karaoke videoRecording, composition, lyric reproduction and sync rightsUse owned or specifically licensed material
Commercial backing-track saleComposition, arrangement, reproduction, distribution and brandingObtain written licences and professional advice
Cover performanceVenue/platform rules and territory-specific licencesConfirm the rules before recording or broadcasting

Practical rule: create karaoke tracks from music you wrote, recorded and control; material commissioned with sufficient written rights; public-domain material using a recording you also control; or music specifically licensed for the intended karaoke use. For a public, monetised or commercial project, obtain advice for the actual jurisdiction and distribution plan.

General information only: this section is not legal advice and does not determine whether a specific use is permitted.

11 Fix the failure, not the whole file

Why your karaoke track sounds wrong.

Name the audible problem first. Each failure has a different repair, and broad processing often creates a second problem.

01 Lead words remain

Wide doubles, ad-libs or effects were grouped with the music. Try a multi-stem split, then automate only the exposed phrases.

02 Backing vocals vanish

The system correctly treated them as vocals. Use the isolated vocal stem selectively, or accept a lead-free version without the original harmonies.

03 Drums lost impact

Snare and cymbal energy overlapped the vocal. Layer the drum estimate from a multi-stem job or replace a few damaged hits.

04 Backing sounds watery

Lossy encoding, reverb and dense overlap created unstable detail. Return to a better source and avoid aggressive denoising.

05 Chorus feels empty

Large vocal stacks occupied much of the arrangement. Rebuild from individual stems, or re-record missing harmony parts.

06 Key change sounds artificial

The transpose is too large or the algorithm is exposing artifacts. Reduce the shift, test formant controls or re-record the backing.

07 Lyrics feel late

The line appears at the first syllable instead of before it, or the final video encode drifted. Advance the cue and check the rendered file.

08 Live playback clips

The backing was mastered too hot for the playback chain. Lower the file level, leave headroom and test with the microphone active.

When to stop repairing and choose another route

Stop when every repair removes more music than vocal. Try a different source file, a multi-stem model, the exact official instrumental, or a fresh re-recording. A busy party track may tolerate light residue; an exposed piano ballad for a professional vocalist may not. The destination decides the threshold.

Need control over the damaged parts?

Split vocals, drums, bass, guitar, piano and other instruments, then rebuild a backing track around the strongest estimates.

12 Direct answers and evidence

Karaoke track FAQ.

Short answers to the questions people ask before turning a song into karaoke.

What is the easiest way to make a karaoke track from a song?

Upload the cleanest authorised copy to an AI vocal remover, download the instrumental, preview difficult sections, repair obvious vocal bleed, then add lyrics or performance cues if needed.

Can I make a karaoke version of any song?

You can attempt separation on most finished songs, but no method guarantees a clean result from every mix. You must also have the rights required for your intended private, public or commercial use.

Can I make a karaoke track for free?

BTR lets you test vocal separation in a browser without installing production software. Current limits and access options are shown on the Vocal Remover page.

What is the best file format for making karaoke?

Use a genuine WAV or FLAC source when available. A high-quality MP3 or M4A can work, but converting a compressed file to WAV does not restore discarded detail.

Why can I still hear vocals in the instrumental?

Backing vocals, stereo doubles, reverb and delay can overlap the music and remain in the accompaniment estimate. Try multi-stem separation and automate only the exposed phrases.

How do I keep backing vocals but remove the lead singer?

Separate the vocal stem, then selectively return backing-vocal phrases with automation. Perfect lead/backing separation is not guaranteed because both are voices and may overlap.

Can I make a karaoke track on an iPhone or Android phone?

Yes. Use a modern browser to upload, process, preview and download the instrumental. A computer is more convenient for detailed repairs, key changes and synced lyric video.

What is the difference between a vocal remover and a stem splitter?

A vocal remover creates vocals and instrumental. A stem splitter can also separate drums, bass, guitar, piano and other instruments for deeper control over the backing.

How do I change the karaoke track to my key?

Find the original key, test the highest and lowest phrases with the singer, then transpose in semitone steps. Render the chosen key before final lyric timing.

How do I add lyrics to a karaoke song?

Prepare authorised and accurate lyric text, split it into singable lines, mark phrase entrances using the isolated vocal, then display each line before the singer needs it. Check synchronisation after final video export.

Is it legal to make karaoke tracks from copyrighted songs?

Creating a file does not grant rights to publish, sell, perform or synchronise the recording, composition or lyrics. Rules vary by use and jurisdiction; use material you control or have licensed and get advice for public or commercial use.

What should I export for live karaoke?

Keep a lossless WAV master and create the format required by the playback system. Leave headroom, test with the microphone and PA, confirm the ending, and keep an offline backup.

Evidence and primary sources

  1. Rafii et al., MUSDB18-HQ audio source separation dataset — 150 full-length 44.1 kHz stereo WAV songs, 100 training and 50 test, with mixture, drums, bass, vocals and other sources.
  2. Stöter et al., Open-Unmix, Journal of Open Source Software (2019) — open-source reference implementation for music source separation.
  3. Hennequin et al., Spleeter, Journal of Open Source Software (2020) — fast music source separation with pre-trained models.
  4. Rouard, Massa and Défossez, Hybrid Transformers for Music Source Separation (2022) — HT Demucs architecture and reported 9.20 dB SDR with extra training data.
  5. Lu et al., Music Source Separation with Band-Split RoPE Transformer (2023) — reported MUSDB18-HQ and SDX23 results for BS-RoFormer systems.
  6. Défossez, Hybrid Spectrogram and Waveform Source Separation (2021) — objective and human-listening results for hybrid separation.
  7. Sutcliffe et al., Evaluating Music Source Separation: Theoretical and Practical Considerations (2025 preprint) — listener ratings and source-dependent metric findings.
  8. Manilow et al., Looking Beyond Averaged Metrics in Music Source Separation, ISMIR 2025 — analysis of distribution-level and perceptual evaluation.
  9. Apple iPhone User Guide: Sing along with Apple Music and Apple Music Sing announcement — adjustable vocal level and real-time lyrics.
  10. U.S. Copyright Office: Musical compositions and sound recordings — explains the two distinct works.
  11. U.S. Copyright Office Fair Use Index and Fair Use FAQ — case-specific four-factor analysis and no fixed safe percentage.
  12. BeatsToRapOn AI Vocal Remover and AI Stem Splitter — current supported workflows, formats and output options.

Continue building