I'm Ilias, the founder of Orcah Studio. I'm building a local-first desktop app (Mac OS available now) that can transcribe audio, recognize faces, detect objects, perform OCR, and describe video scenes.
Plus, a local-first video agent that understands your prompt and finds the exact moment you're looking for. I fine-tune the agent based on an open-weight model (an 8-billion-parameter model) with tool calling.
You can check out all the features here: https://orcah.app.
Plus, I have an MCP that you can connect to Claude Code, Codex, and OpenCode to search your indexed videos: https://orcah.app/mcp
And transcription with speaker diarization: https://orcah.app/features/transcription
On a more serious note, I think on-device object/face recognition search is a great idea. Having the cta button on the site go to Polar is a little unexpected. It'd be nice to be able to try the app without going through the hoops of running it in a Docker container, if that's the same app or system?
Noted. Would like to show the app in a 15-minute call here (https://cal.com/iliashadad/orcah-studio-demo)
Yes, the Docker container is pretty different than the app, but it does share similar infrastructure with richer video scene data, direct integrations with Davinci Resolve, Final Cut Pro, and Adobe Premiere Pro, and a smart video search agent.
Thank you for your comment, much appreciate it