Back to blog

P(doom) Video in 2026: How a Music Video Emerges from Code

Hello HaWkers, the open source pdoom-video project turns a song into a music video whose images are calculated from the track's timeline. According to its repository, the preview runs in a browser, while export uses the same logic to produce 1080p video at 60 frames per second, with a 4K rendering option. It is a concrete example of generative art in which code serves the visual direction, lyrics, and rhythm.

What can you learn from a project like this without copying its aesthetic or its music? In this article, we will separate creative decisions from technical ones, recreate a small version of the core idea, and explore the limits of performance and copyright. The result is a way to think about your own music video, audio visualizer, or interactive piece.

What makes this different from a manually edited video?

In a traditional editor, you place elements on a timeline and save the result. Here, every frame is calculated from a question: what should appear at time t in the song? The project's engine documentation describes the scene as a deterministic function of time. If you request the same instant again, the renderer should produce the same composition. That lets you inspect a single frame, return to a specific word, and export without depending on the path playback took to get there.

The project uses TypeScript, three.js, and a web application for previews. The song file, word timings, and audio analysis are inputs; scenes live in separate modules. During export, a headless browser produces frames and ffmpeg assembles the final file. This division matters because a creator can review composition, movement, and synchronization in stages. The finished video does not materialize from one command: visual direction and corrections remain editorial work.

There is also an important distinction between video generated by code and video generated by an image model. The README credits a collaboration with Claude Code in building the project, but the final frames are drawn by the renderer written for this work. AI took part in the process of creating the software; it does not automatically replace choices about framing, pacing, and review. The result can be reproduced, inspected, and changed scene by scene.

Music, words, and time are the real inputs

Synchronization does not begin with a flashy transition. It begins by deciding which musical events control the image. The repository contains word-level lyric timing, along with beats, sections, instrument onsets, and intensity measurements. Its documentation mentions source separation, forced lyric alignment, and a speech-recognition check while preparing those data. You do not need to repeat that analysis to run the renderer with the already versioned data.

Keeping these concerns separate avoids a common mistake: hard-coding the appearance of a line at a second chosen by ear, then discovering after an edit that it no longer matches the voice. When one scene is anchored to a word and another to a beat, the intent becomes explicit. You can move a scene and still check which verse or musical event justifies it. The same applies to karaoke typography: every word needs its own start and end if the highlight is to follow the performance rather than the entire line.

The author says in the README that the analyzed track is approximately 132 beats per minute. That number belongs to the song used in this project; it is not a rule for other videos. Before borrowing the approach, obtain data for your own soundtrack and visually verify the edit points. A technically correct beat marker can still fall at the wrong place for the meaning of a lyric. Listen, watch, adjust, and compare again.

How to try the preview without rendering the whole video

The shortest path is to open the project's own application. The README lists Bun for dependency installation and Vite for starting the preview; Chrome and ffmpeg enter the export workflow. If you want to study the visual direction, the preview is enough: you can pause, move frame by frame, and switch scenes before spending time encoding a final file.

# Clone the open source code and enter the music video application.
git clone https://github.com/mexicat/pdoom-video.git
cd pdoom-video/app
bun install
bunx vite

The preview opens at http://localhost:5173; according to the documentation, the ?t=23 parameter starts it at a specific moment. Space pauses, arrow keys move forward or backward, and comma and period let you examine frames. These shortcuts help answer a precise creative question: does the image convey the idea when the music reaches that word?

For a reliable review, select three moments: the start of a phrase, its peak of energy, and the transition to the next scene. Compare type size, contrast, and legibility at each point. A composition can look excellent when still and become unreadable once motion starts. A quick preview is where you catch that problem before an expensive render.

A small experiment with deterministic frames

You do not have to reproduce the entire pdoom-video architecture to test the concept. A canvas and a function that takes time are enough to draw an element that responds to music. The example below creates a pulsing bar and displays a word only within a chosen interval. The timing numbers are demonstration values; they do not represent the lyrics or timestamps of the original project.

<canvas id="palco" width="960" height="540"></canvas>
<script>
  const canvas = document.querySelector('#palco');
  const ctx = canvas.getContext('2d');

  function quadro(tempo) {
    // The same time input always produces the same image.
    const pulso = (Math.sin(tempo * 6) + 1) / 2;
    const mostrarPalavra = tempo >= 2 && tempo < 3.2;
    ctx.fillStyle = '#101820';
    ctx.fillRect(0, 0, canvas.width, canvas.height);
    ctx.fillStyle = '#f3b54a';
    ctx.fillRect(100, 390, 760 * pulso, 18);
    if (mostrarPalavra) {
      ctx.font = 'bold 72px sans-serif';
      ctx.fillText('CREATE', 100, 270);
    }
  }

  quadro(2.5); // Change the time to inspect another frame.
</script>

The example is intentionally simple. In a real work, the width of the bar could come from the track's measured intensity, and the word interval could be read from an alignment file. The principle remains: quadro(t) depends on t and on known data, not on how many times the animation has run. That makes it possible to return to a scene and revise a detail without waiting for the whole earlier sequence.

If you introduce particles or random texture, use a stable seed derived from time or from the identity of an element. Calling Math.random() on every render can produce a different result when you request the same frame. The project's guide explicitly recommends seeded randomness for deterministic scenes. This choice becomes even more important when export calculates multiple samples of a frame to simulate motion blur.

From preview to export: why does the cost rise so much?

Export is not the same as recording the browser screen. The repository's script asks the browser for frames at defined instants and passes that sequence to ffmpeg. In the default described by the README, the finished video is 1920 × 1080 pixels at 60 frames per second, with x264 video and AAC audio. A 4K mode renders layers at the larger resolution; it does not merely enlarge an already finished image.

# Generate a short version to review movement and cuts.
cd pdoom-video/app
bun scripts/render.ts video --from 20 --to 25 --out ../out/test.mp4 --preset veryfast

# After reviewing it, export the complete music video.
bun scripts/render.ts video --samples auto --shutter 0.2 --out ../out/music-video.mp4

The first command follows the engine guide's short-test example. The second uses the complete workflow documented in the README. Before running them, check that Chrome and ffmpeg are available and that the output directory exists. A short test reveals stalls, abrupt cuts, and typography problems sooner than a complete render. It also makes it easier to compare two versions of a scene without turning every adjustment into a long wait.

Motion blur through sampling changes the cost most. Rather than calculating just one image for each frame, the renderer combines images from nearby instants. The README explains that automatic selection uses more samples during fast movement. This improves the continuity of a zoom or abrupt movement, but it multiplies GPU work. The author describes up to 324 subframes in fast passages; that is a limit in the documented workflow, not a performance promise for every computer.

Encoding adds another cost. Film grain and fine detail are hard to compress, especially at 4K. If the goal is web publication, compare visual quality, file size, and export time. A review version can use fewer samples; the final version only needs to pay for extra quality where the difference is visible. Testing a demanding passage before the whole video is a practical decision, not an artistic compromise.

AI helps write scenes, but someone still has to direct them

The README says that the concept, treatment, audio analysis, renderer, and scenes were developed in conversation with Claude. That also makes the project interesting to people following AI tools. Its verifiable contribution, though, is the open process: it provides documentation of the visual language, scene modules, synchronization data, and preview commands. We can study how decisions became editable artifacts.

A useful workflow for your own projects is to write a short intention for each passage first: what emotion it should convey, which word gets emphasis, and what should change in the visual rhythm. Then turn that intention into a small scene and request a still image of its key moment. Move on to animation and export only when the composition is legible. An AI tool can suggest alternate code or treatments, but a review needs to compare the result with the intention. The measure is not how many lines were generated; it is whether the scene works with the song.

This practice connects with our article about generative animations as a meeting point between art and programming. The language and tool can change, but the creative question is similar: which rules produce an expressive image, and how do you know that image deserves to stay? With pdoom-video, the answer also depends on musical time and on how each verse is read.

Open source code does not automatically free the music

There is an easy trap when you find a visual project in a public repository: assuming that everything inside it can be reused in the same way. The README says the code is distributed under the MIT license, while fonts retain their own licenses and the music and lyrics belong to their respective creators. That matters both when publishing a modified version and when using excerpts in a commercial demonstration or on social media.

If you want to learn from the project, study its architecture and create scenes using assets you own or are authorized to use. If you want to reuse audio, lyrics, typography, or renders, check each material's specific license and obtain any needed permission. Authorship of a renderer and authorship of a musical work are separate questions. Crediting a source is good editorial practice, but credit alone does not replace permission when it is required.

It is also worth distinguishing taking inspiration from the method from copying the visual identity. The project's documentation describes color, composition, and transition choices made for that particular song. For your own video, begin with the meanings and rhythms of your own track. If your first decision is to recreate the example's most striking effect, the result may look like a technical demonstration unrelated to the music.

What to take into your next creative project

Pdoom-video shows that a code-generated music video can be both a visual work and a reproducible system. The central lesson is not a specific library: give every scene an intention, connect visual events to verifiable audio data, and keep preview and export close to each other. That lets you correct a late word, revise a frame, and compare versions without losing the work inside an opaque sequence of manual steps.

Start small: select a short piece of music you are allowed to use, mark one word and one beat, draw a deterministic frame, and export just a few seconds. Then watch the video with sound and ask whether the motion follows the meaning of the track. The next improvement might be typography, depth, blur, or a new scene. The real gain comes when each feature solves an identifiable expressive problem.

Projects like this also leave a question for the future of digital creation: if writing scenes gets faster with AI assistance, where do we focus our attention? Probably on choosing what to show, judging the result, and taking responsibility for the material used. Those parts remain visible to the audience even when the code stays behind the scenes.

Let's go! 🦅

📚 Want to Keep Up With What Is Coming?

This article covered making a music video with code, but the ecosystem changes every week, and not every experiment becomes an article here.

On X, I share what I am testing, behind-the-scenes notes from projects, and new developments before they become posts.

Follow Me There

👉 Follow @jeffbruchado on X

💡 Daily content about development, careers, and the tools I actually use

Comments (0)

This article has no comments yet 😢. Be the first! 🚀🦅

Add comments