Over two days, I took Cutwise from an empty desktop project to a working Mac application that can transcribe a recording, suggest edits, add polish and render a finished video.
It can now process a real screen recording from import to export. The Mac build is signed with a Developer ID certificate, notarised by Apple and accepted by Gatekeeper. That sounds like a finished product, but it is not one yet. Windows still needs native testing, the Mac build needs a genuinely separate clean-machine test, and several distribution and legal checks remain open.
The interesting part of the build was not how quickly the screens appeared. It was how often the first apparently successful result exposed the next problem.
Starting with three timelines
I began with one architectural rule: Cutwise should never alter the original recording.
The project separates the source recording from the decisions made about it. The source timeline holds the video, audio and transcript. The edit timeline records cuts, shortened pauses and transitions. A separate overlay timeline holds callouts, titles and captions.
That structure made the first version less glamorous but much safer. A project stores references, timestamps and immutable snapshots rather than repeatedly rewriting media. Accepted edits can be reviewed later, old decisions remain available for recovery and the original file is checked before and after important operations.
I built the desktop foundation in Flutter for macOS and Windows. The first phase covered project creation, drag-and-drop import, metadata, embedded playback, autosaving and secure provider-key storage. FFprobe reads the source details, while API keys stay in Keychain or the Windows equivalent instead of entering a project file.
Even this small foundation found a useful failure. Native playback stalled during one early release test. I did not find a conclusive underlying cause, so I added a bounded startup deadline and a retry path rather than allowing the interface to spin forever.
Making transcription part of the editor
The second step was local transcription. Cutwise extracts audio with FFmpeg and sends it to a private Python worker running Whisper. The worker processes bounded windows, returns source-relative word timings and shuts down after the job. If acceleration fails, it can retry on the CPU without keeping a half-finished transcript.
Model downloads are explicit and resumable. Transcription works without a cloud account, and failed or cancelled jobs leave the previous result active. The transcript is not just text beside a video: clicking a word seeks the source, while playback highlights the current segment and word.
Keeping every timestamp in the original recording's clock became important almost immediately. Silence, filler words, semantic suggestions and later visual observations all needed to refer to the same stable source. Converting too early into an edited timeline would have made every later decision depend on whichever cuts happened to be active at the time.
The first cleanup layer therefore became an edit decision list rather than a destructive operation. It proposes long pauses and a small set of transcript-visible filler words, then lets me preview, accept, reject or replace each range manually.
Rendering those decisions revealed a less obvious problem. A test containing 118 cuts exceeded FFmpeg's expression parser depth when the selection expression was built as one flat chain. I changed the renderer to build balanced expressions and kept the long fixture as a regression test. The final two-minute timing check preserved all 120 audio and video markers without progressive drift.
Letting AI suggest without letting it decide
Basic silence removal could not tell whether somebody had restarted a sentence, repeated a demonstration or corrected an earlier explanation. I added an optional semantic layer for that work.
Cutwise supports OpenAI, Anthropic and Gemini through a shared provider boundary. The user supplies the key and chooses the model. Transcript text is treated as untrusted media, responses have a strict schema and every proposal must point back to known words and source ranges.
Most importantly, suggestions remain suggestions. Semantic edits never qualify for bulk acceptance. The review shows what would be removed, what would be kept and why the model proposed the change. A failed or cancelled analysis cannot replace an existing edit plan.
I applied the same rule to the polish features. Crossfades and dips live in their own transition layer instead of changing the source cuts. Ordinary speech cleanup stays as a hard cut, with only a tiny audio-edge fade to avoid clicks. Text callouts are anchored to source words or explicit output times, use bundled fonts and become pixels before they reach FFmpeg, so generated text cannot become part of a filter command.
Visual analysis is optional and deliberately sparse. Cutwise samples a bounded set of screenshots locally, reduces obvious repetition and lets the user inspect exactly which frames would be sent to a provider. The complete video is never uploaded. Visual observations can support an edit, but a static screen alone is not treated as permission to remove narrated material.
By the end of the feature work, the application also had captions, audio loudness correction, gentle noise reduction and export presets. Each arrived as another derived layer over the same original recording rather than a reason to rebuild the editor around a conventional multitrack timeline.
The real recording disagreed with the fixtures
Synthetic media let me verify timestamps, rendering and failure handling, but it could not tell me whether Cutwise made sensible editorial decisions.
I ran the installed app through its complete workflow using a 209-second human screen recording and a configured OpenAI model. Cutwise transcribed 527 words, proposed silence and semantic edits, sampled the screen, generated callouts, produced captions and rendered two review exports.
The results were useful precisely because they were imperfect.
One semantic suggestion marked an intentional topic heading as a false start with 0.96 confidence. I rejected it. Two other suggestions were plausible enough to accept for the listening draft. Visual analysis described a largely static narrated screen as possible dead time because a request had not received enough transcript context from across the full interval. I changed the context selection and added regression coverage, but I did not pretend the earlier result had become correct.
Only two of eight proposed callouts survived review. Some merely repeated the narration; others were inaccurate or described a fixed problem as though it still existed. The accepted cards still needed their placement changed to avoid covering the presenter or the page content.
Caption review exposed two deterministic problems. Version punctuation gained unwanted spaces, and greedy phrase splitting could leave a single word flashing on its own. Both fixes started with failing tests and kept the original transcript and word timings unchanged.
That validation established the rule I now trust most in the project: a high confidence number is not an editorial decision.
Finding the filler words Whisper had removed
The most revealing failure was that Cutwise found no filler words in the real recording.
The cleanup code was working as designed, but the main Whisper transcript had tidied the speech while transcribing it. The three hesitations I could hear did not exist as transcript words, so there was nothing for the original matcher to inspect.
Changing the main transcription prompt recovered some fillers but also created a repetitive um hallucination. Rather than making the visible transcript less trustworthy, I added a separate local acoustic pass used only for filler evidence.
That pass analyses overlapping 20-second windows, confirms each candidate again in a smaller window and checks the surrounding audio for voiced sound and safe boundaries. It does not rewrite the transcript or automatically accept a cut. Uncertain cases remain observations that can be auditioned and converted into manual edits.
The new pass found all three labelled locations in the human recording. I listened to the Original and Proposed comparisons before applying them. On the checked-in speech fixture it found all nine constructed hesitations on both Apple Silicon and CPU paths without detections outside the labelled ranges, although that still does not make it a universal filler classifier.
Discovering that packaging is part of the product
The editing pipeline was only half of the work. A normal user should not need Python, FFmpeg, developer tools or a carefully prepared model cache.
I built a private runtime containing Python, the transcription worker, FFmpeg, FFprobe and the required native libraries. Models remain separate downloads so they can be selected and removed without rebuilding the app. The packaging work also grew its own checks for licences, source materials, native-library ownership, runtime hashes and the exact bytes included in an installer.
This was where the development machine became actively misleading.
A fresh-user simulation downloaded a complete model snapshot, then failed to load it offline. The worker had downloaded a pinned revision but later asked the cache for main. My old development cache already contained that reference, so every earlier test had quietly hidden the bug.
I changed the worker to record the exact verified revision only after a complete download. Cancelled or incomplete downloads cannot promote that reference. I rebuilt the runtime, signed the app again and sent both the app and DMG through a new notarisation cycle.
The current Mac candidate is version 0.9.2. Its app and DMG are Developer ID signed, notarised, stapled and accepted by Gatekeeper. The latest maintenance pass also gave provider keys a recognisable Cutwise label in Keychain, reduced duplicate progress displays and made errors much harder to miss. The full Flutter suite now contains 444 passing tests, alongside worker, packaging and native credential checks.
Where Cutwise stands after two days
Cutwise can now import a recording, transcribe it locally, suggest silence, filler and semantic edits, analyse selected screenshots, review transitions and callouts, generate captions, improve audio and render an MP4. Projects reopen with their decisions and export history intact, and the original media remains untouched.
The Mac application has crossed the technical signing and notarisation gates, but it remains marked INTERNAL ONLY. A simulated fresh-user environment is not the same as installing it on a separate Mac with no development history. Windows packaging exists but has not been validated on a Windows host. Application terms, private-beta terms and some combined-distribution dependency reviews also remain unresolved.
So the next step is not another editing feature. It is to validate the Windows build, repeat the Mac workflow on genuinely separate hardware and put the current application in front of a small group of testers.
Two days produced far more than the first desktop shell, but they also made the unfinished work much clearer. That is a better milestone than pretending the notarised DMG means Cutwise is ready for everyone.
