Edit: I'm curious which LLM was used to generate the code. I fed the title of your post to claude/deepseek/qwen/codex asking to recommend a stack for this project, expecting to frown thinking that they still recommend tesseract. However, I found that they all recommend apple's vision framework. In fact the latest model to recommend Tesseract is gpt-4.1.
Later that evening, I just paused the YouTube video on my phone, circled the part of the display with Google Lens, and it read the whole thing with pretty good accuracy. It was mindblowing :)
A hand coded electron project would have a discussion about electron vs native. Relevant in general but off topic in the context of this particular app.
What’s the downside risk of having "worse software" when you’re just ideating and putting things out there to see how people like it.
You're totally right. Apple Vision is generally way faster and more accurate on Apple Silicon. Pure Mac-only, I'd use it. (also smaller)
Main reason I went Tesseract: I want SCM to stay portable easily. It's ARM Mac for now, but the inference layer is all JS end-to-end — Transformers.js + ONNX for CLIP/SigLIP + Whisper, Tesseract.js WASM for OCR, all in plain Node workers.
A bit of background: I'm actually an iOS/macOS dev and I really love SwiftUI and AppKit — I just wanted v1 to stay portable by construction. Exploring a native Swift + MLX v2 track separately for speed.
Can you copyright things like this now that LLMs exist? I mean, up until now if a small startup has a great idea they will get bought out by big tech which will integrate (or kill) their tech. But now with LLMs can the likes of OpenAI just tell their model to make something that works similar to X (such as this project) and then get round copying laws and negate being behind the curve?
EDIT: switched to the correct spelling of copyright.
No one thinks Apple violated Sherlock copyright. They just made it (nearly) functionally redundant.
(& by "this" I mean an approximate AI search for photos & videos - I can't account for the "every frame", nor for the comparative search quality)
Just for product aesthetics, i believe it created a thumbnail folder with every resolution, which is annoying since i have a multi-decade 8tb library.
There were other actual issues with organization and display. But with this I realized I can just get gemini to write a DAM for me in a weekend.
The view options are what killed it in the end.
The thumbnails issue was just me looking at the sausage and like "ugh, I wouldn't have done that."
It's always a tradeoff between time and work. I would have done a sliding window around the current screen, but then you might have placeholders. Maybe it should do low-res first and generated hi-res thumbnails 3 pages out while scrolling. But that's a lot of work for the 10% case.
What I want is aperture back, with all the fun AI stuff and using Affinity for photo editing (which I have anyway). I'll add that to my queue of projects I guess.
And then you end up with image metadata processing code like this, which just by a cursory glance I’m sure has edge case bugs : https://github.com/allenv0/SCM/blob/main/screenshot-probe.js
Based on my experience, sampling rate can be tricky if what you are looking for lasted less than interval period.
you are pirating first-release movies for commercial purposes?
isnt it expensive tho?
Tall order but would love to get my cpu back!
otherwise, I can just make my own with my own ai. why consume someones slop when I can eat my own.