Overview
Find is a local-first image intelligence platform: it uploads, indexes, searches and clusters images on your own machine. Image processing, vector generation and search all stay inside your local stack, so a personal photo library never has to leave it.
- Upload individual images or ZIP archives.
- Extract captions, detected objects, OCR text, EXIF metadata and dimensions.
- Search with natural language over hybrid embeddings.
- Cluster related images automatically once indexing completes.
- Share albums with scoped links, and keep hidden images in a password-gated vault.
Architecture

The frontend is Next.js with React Query. The backend is FastAPI with SQLAlchemy, PostgreSQL and pgvector, Redis with RQ workers, and MinIO for object storage. The ML pipeline runs YOLO for detection, BLIP for captions, PaddleOCR for text, SigLIP for embeddings, InsightFace for faces and HDBSCAN for clustering.
It ships as Docker Compose profiles, so the same app runs with no AI at all, with deterministic mock outputs for tests, on CPU, or on an NVIDIA GPU.
Results
- images processed
- 100k+
- recognition accuracy
- 98%
- query latency
- −350 ms
- index size
- −20%
| Metric | Before | After | Change |
|---|---|---|---|
| Image throughput | 100 | 130 | +30% |
| Memory usage | 100 | 85 | −15% |
| Index size | 100 | 80 | −20% |
Vector quantization and compression cut the index by 20%, which is what made 100,000+ images practical on one machine. Resolving three pipeline bottlenecks took 350 ms off query latency.
Screens



Open source
Find is AGPL-3.0 licensed and open to contributors through GirlScript Summer of Code 2026, with issues labelled by difficulty so first-time contributors know where to start.
