Developer Allan Lee (allenv0) published the open-source project SCM (Screen Memories) under the MIT license on Show HN — local AI search for all photos and every video frame in any folder on macOS, without accounts, cloud, or uploads. All inference is performed on the Mac itself, and the app is installed on Apple Silicon with three Homebrew commands. Search covers five modes: from semantic search of frames using swappable CLIP and SigLIP models to jumping to an exact timestamp, text in images, and spoken lines.
What happened
In the Show HN section on Hacker News, developer Allan Lee (allenv0) presented SCM (Screen Memories) version 0.2.4 — a local AI search utility for media files in any folder on macOS, distributed under the MIT license. Search works in five modes. Semantic search of files relies on four swappable models via ONNX Runtime: by default CLIP ViT-L/14@336 with weights of about 435 MB and a latency of approximately 480–570 ms per frame on CPU, with alternatives being SigLIP-2-B/16, SigLIP-2-L/16@256, and SigLIP-B/16@384. Scene search within videos uses shot boundary detection via ffmpeg with presets from Eco with a point every 60 seconds to Ultra Pro with a 2.5-second step and a jump to the exact timestamp. OCR is implemented via Tesseract with English plus 35 switchable languages, and search for spoken lines is via Whisper with tiny.en models of about 150 MB or base.en of about 300 MB. The optional Ask mode opens a local chat via a llama.cpp-sidecar, available only on loopback, with Qwen3 1.7B of about 1.1 GB by default or Llama 3.2 3B. Installation on Apple Silicon with macOS 12+ is described in three Homebrew commands, with the key steps being brew tap allenv0/scm and brew install --cask allenv0/scm/scm.
Context
The engineering foundation consists of long-known open components: CLIP and SigLIP provide frame embeddings, ffmpeg segments videos into shots, Whisper transcribes speech, Tesseract recognizes text in images, and llama.cpp handles local chat. The novelty here is not model-based, but system-engineering: the author assembled these parts into a single media search pipeline and published the full code along with honest measurements — latency per frame, weight size, and segmentation budgets. In terms of capabilities, this is an open-source analog of Apple Photos' built-in search and proprietary utilities, but with a level of detail of "a specific frame or line" that is usually not present in standard solutions. The design of the Ask mode is also noteworthy: it works as RAG over already extracted text metadata — lines, OCR text, and file names — with numbered citations and tok/s streaming, rather than as multimodal understanding of frames.
Why this matters for the industry
For the industry, this is a signal of commoditization: the layer of semantic search for personal photos and videos is already being assembled entirely from open components on a consumer Mac, without cloud APIs or subscriptions. Selling such search on macOS itself has become harder, while assembling it as a feature in your own product has become faster and cheaper. For builders, SCM is useful primarily as a proven MIT-licensed template: a multimodal index of embeddings, shots, Whisper transcripts, and OCR plus a local LLM-sidecar — the code and patterns can be transferred to your own products. Defending in this category now requires vertical workflows, indexing speed, and integrations, rather than the embeddings themselves. There is no direct market shift today — the project is at version 0.2.4 without notable traction — but the benchmark for the category has already shifted.
Why this matters for users
For a reader with a Mac on Apple Silicon, the project is available today: model weights are downloaded once, and everything works offline thereafter, and media does not go anywhere. Search is performed by a natural language description of a frame in five modes: you can find a scene in a video and jump exactly to the right timestamp or line, find text in images via OCR, and ask questions to your library in Ask mode with citations. The repository has ready-made installation commands and a demo GIF, so you can check out the project immediately after reading.
What is still unknown / limitations
There are no public search quality metrics yet: no recall, no comparison with Apple Photos, and at the time of publication the project had zero traction on Hacker News. Initial indexing of a large archive is a non-trivial task: with the arithmetic from the published latency of about 0.5 seconds per frame, indexing 10,000 photos will take about an hour and a half of pure CPU, and for video the time will increase even more noticeably — this is an estimate, not a measured benchmark. Line search works only in English, since the Whisper models included are tiny.en and base.en, and the multilingualism of the semantic mode is limited by the capabilities of CLIP and SigLIP themselves and requires verification for the Russian language. Predictions about the emergence of a wave of similar builds and the transformation of local multimodal indexes into a standard layer of file managers and OSes are interpretations, not facts.
Sources
- GitHub repository allenv0/SCM (Screen Memories project)
- Show HN: AI search for every photo and every frame of video on macOS — discussion on Hacker News
Author
Look at AI, editorial team
