[Case Study] Completing a Desktop App in 1.5 Hours by "Chatting" with Codex: The Mind-Blowing Impact of VibeCoding
Stop "writing code." Start "defining completion." This is a real-world experimental log validating the AI-driven development methodology known as "VibeCoding." Moving straight from a PRD, a cross-platform desktop application (MediaSplitter AI) featuring video splitting and AI transcription was fully built in just 1.5 hours. This article opens up the entire blueprint for development process innovation in the AI era, alongside concrete, actionable know-how for success.
Stop "writing code." Start "defining completion."
Recently, VibeCoding has become the talk of the town among software engineers. On top of that, OpenAI's Codex has evolved phenomenally since April, firmly securing its spot as the absolute king of AI development tools. By combining these two, I ran an experiment to bypass the heavy, traditional "Design $\rightarrow$ Implement $\rightarrow$ Test" pipeline, building a product at breakneck speed simply by aligning my "vibe" with the AI.
Today, I’m opening up the complete log of developing the media processing tool "MediaSplitter AI" via VibeCoding. From handing over the PRD (Product Requirement Document) and a hand-drawn UI mockup to compiling a desktop application with built-in AI (Whisper), Codex essentially completed the entire application in just about 1.5 hours👍.
1. Project Overview: What is MediaSplitter AI?
It’s a tool designed to solve those everyday frictions where opening full-blown video editing software feels like overkill—such as when you just want to "trim from point A to point B" and "maybe transcribe the audio while you're at it."
- Core Features: High-speed video splitting (FFmpeg), AI transcription (Whisper), audio extraction.
- Tech Stack: Tauri (Rust) + SvelteKit + Tailwind CSS.
2. Development Log: The Breakneck Timeline
STEP 1: Generating the MVP straight from the PRD Codex time: 8 minutes
First, I threw the Markdown PRD and the mockup image at the AI.
In just 8 minutes, the AI spun up a web-based prototype featuring the following:
- Deliverables: A pro-grade dark mode UI, preview playback, and stream copying via FFmpeg (blazing-fast processing with zero re-encoding).
- The Vibe Highlight: In response to my prompt to "just make something that works first," it instantly suggested a Node.js backend and frontend architecture.
STEP 2: Chasing a "Sleek" UI/UX Codex time: 10 minutes
Next, we tweaked elements tied directly to user experience, like adding buttons to set the Start and End markers or visually highlighting the selected range.
- Added Features: Timeline indicator lines and selection highlighting.
- The Takeaway: In VibeCoding, the AI successfully translates vague requests like "make it feel a bit more like this" into code, making the micro-adjustment feedback loop incredibly fast.
STEP 3: Going Desktop with Tauri Codex time: 15 minutes
We took the polished web app and migrated it into a desktop application (Tauri) that runs natively on Mac and Windows.
- Technical Breakthrough: Shifted direct local file operations—which are notoriously difficult in a browser environment—to Rust backend processing. It even auto-generated the GitHub Actions configurations for building binaries.
STEP 4: Integrating AI Transcription & Solving Portability Human & Codex dialogue: 1 hour
This was the most dramatic phase. Initially, we used the Python version of Whisper, but the environmental dependencies (forcing users to install Python and various libraries) was a massive dealbreaker for distribution.
- The Solution: The AI proposed switching to
whisper.cpp(Rust bindings). - Final Form: Successfully configured the app to bundle the FFmpeg binaries and the pre-trained Whisper models right inside the app package.
- The Result: A true portable application was born—"Just install it, and you can run AI transcription on any PC offline."
3. 3 Keys to VibeCoding Success
Through this project, I uncovered a few golden rules for AI-driven development:
① Provide a Solid PRD Requirement Definition
AI isn't magic. By handing over a clear document (PRD.md) outlining "what you want to build" and "who the target audience is" at the very beginning, the accuracy of the AI's output (the Vibe) improves exponentially.
② Embrace "Disposable Code"
As seen when I asked the AI midway through to "just make a pure HTML version to test," the cost of experimenting with different architectures is near zero. Don't fixate on a single "correct" answer; maintain an "orchestrator" mindset, let it generate multiple patterns, and pick the best one.
③ Solve Setup Bottlenecks Together with AI
Environment and configuration issues—like the Python errors we hit—are typically what developers hate dealing with the most. However, if you paste the raw error logs directly, the AI will set up virtual environments (venv) or offer pivot strategies (like migrating to the C++ version) for you.
4. Conclusion: VibeCoding Only Requires "Curiosity"
The era where "you can't build an app unless you can code" is officially over. MediaSplitter AI was completed entirely through a dialogue with Codex, without me writing a single line of complex Rust code or FFmpeg commands by hand.
Total Development Time: 95 minutes
Lines of Code Written by Me: 0 lines (100% generated and adjusted by Codex)
If you have a note on your desktop filled with "ideas I want to build someday," stop waiting. Throw that note into Codex right now. That is the exact moment your VibeCoding journey begins.
Tech Stack Summary Used in this project
- Frontend: HTML5, CSS3 (Tailwind), JavaScript
- Backend: Node.js (Initial), Rust / Tauri (Final)
- AI: OpenAI Whisper (whisper-rs)
- Engine: FFmpeg (Stream Copy)
Author's Soliloquy:
Something that was running inside a browser just an hour ago is now running smoothly on my desktop as a native Mac .app file. This feeling of absolute capability is the ultimate drug of using Codex.
Source code is available here: https://github.com/perfact-tang/mediaeditor
Additionally, the contents of the PRD (Product Requirement Document) are listed below.
🛠️ MediaSplitter AI: Product Requirement Document PRD
1. Project Overview
- Product Name: MediaSplitter AI (Tentative)
- Concept: "Trim, Export, Transcribe" in the fewest steps possible. A lightweight media processing tool compatible with Mac/Win, powered by AI (Whisper) and a high-speed engine (FFmpeg).
- Core Value: Complete quick trimming, audio extraction, and transcription tasks in just a few clicks, freeing users from the friction of opening heavy video editing suites.
2. Functional Requirements
2.1 Input & Media Management
| Feature | Description |
|---|---|
| Multi-Format Support | Support importing video (MP4, MOV, MKV) and audio (MP3, WAV, M4A) files. |
| Import Method | File selection dialog + Drag & Drop support. |
| Metadata Display | Display filename, total duration, resolution, and file size. |
2.2 Editing & Preview Core UI
- High-Performance Player: Variable speed playback from $0.5\times \sim 5.0\times$.
- Waveform Timeline: Visualize audio amplitude to help users easily identify transition points.
- Segment Management: * Allow users to drop a split point at the current playback position via a "+" button or hotkey.
- Display segments in a list format showing Start/End times, with individual delete and micro-adjustment capabilities.
2.3 Splitting & Extraction Logic
- Blazing-Fast Trimming (Stream Copy): Instant clipping without re-encoding, ensuring zero quality loss (powered by FFmpeg).
- Smart Detection (Optional/Advanced): * Silence Detection: Automatically scan quiet sections and place split markers.
- Scene Detection: Detect visual cuts/scene changes (Future Expansion).
2.4 AI Transcription
- Engine: Powered by OpenAI Whisper (Prioritizing local offline execution).
- Output Formats: Export plain text (.txt) and SubRip subtitle (.srt) files.
- Language Support: Multi-language support covering Japanese, Chinese, Korean, English, and auto-detection.
2.5 Output Settings
- Batch Processing: Export all designated segments simultaneously with a single click.
- Multi-Output Toggles: Provide checkboxes for "Video, Audio, Text" per segment to allow simultaneous, independent exports.
3. UI Layout Plan
Adopting the intuitive layout drafted in the original hand-drawn mockup.
| Section | Role / Functionality |
|---|---|
| Left: Drop Zone | File loading area featuring a prominent "+" icon. |
| Middle: Main Canvas | Video preview window, playback controls, and the waveform timeline. |
| Right: Batch Panel | List of registered segments (e.g., 1. 00:31~ ❌) and export format checkboxes. |
| Bottom: Action Bar | Global progress bar and the main "Execute Batch Split" button. |
4. Recommended Tech Stack
- Framework: Tauri (Lightweight, blazingly fast, native Mac/Win support)
- Frontend: SvelteKit + Tailwind CSS * Backend (Rust):
ffmpeg-next: For high-speed video/audio stream copying and trimming.whisper-rs: For running local offline AI transcription.
- Build/Distribution: Automated
.dmg(Mac) and.msi(Win) compilation via GitHub Actions.
5. Roadmap: 3-Phase Development
【Phase 1】MVP Minimum Viable Product
"Get something working first"
- File loading and media playback functionality.
- Manual segment marking $\rightarrow$ Video splitting using FFmpeg stream copy (no re-encoding).
- Successful compilation checks on both Mac and Windows.
【Phase 2】AI & Advanced Feature Integration
"Fulfill all requirements from the mockup"
- Implement the Whisper-powered transcription engine.
- Render visual audio waveforms on the timeline.
- Connect the multi-output checkboxes (Video/Audio/Text) to the export logic.
【Phase 3】Professional Polishing
"Elevate the user experience"
- Integrate smart silence detection.
- Add open-caption burn-in (Hardsub) capability.
- Implement GPU acceleration to maximize processing speeds.
📝 Message to the Developer You
This blueprint strikes a perfect balance between simplicity and rich functionality. By leveraging Tauri and Rust, we will liberate users from the dread of launching bloated editing suites just to make a quick trim!
