How every uploaded image describes itself: our alt text and semantic search bar
A business uploads hundreds of photos to the panel. Most of them are named "IMG_4821.jpg". None of them have alt text. This is a loss for both accessibility and search engines. Three months later,...
A business uploads hundreds of photos to the panel. Most of them are named "IMG_4821.jpg". None of them have alt text. This is a loss for both accessibility and search engines. Three months later, when searching for "that photo in front of the shop", you still have to browse folder after folder.
Every image uploaded to Marcapony gains a description and meaning that can be searched for without anyone doing anything. In this article we explain how that pipeline works.
Upload: "uploaded" vs "ready" are separate things
The browser uploads the file in chunks directly to S3 using pre‑signed URLs generated by the API. The file does not pass through the API server. Once the upload is complete, the file is copied under {hesap}/raw/.
This raw/ prefix is the trigger of the pipeline. A file being uploaded and a file being ready for the service are two separate states and follow two separate paths.
The trigger should be fast, the work can be slow
For every file that lands under raw/, an EventBridge rule sends an event to the trigger endpoint of our media processing service. EventBridge waits for a 5‑second response. Video conversion can take minutes. Therefore the trigger endpoint does not perform any work itself: it decides which processes the file will go through, enqueues a task in Cloud Tasks, and immediately responds.
The real work is done on a Bun service running on Google Cloud Run in the Frankfurt region. The queue handles retries and concurrency on our behalf.
Images and videos
Images are converted to WebP. The compression quality is not fixed; it is chosen progressively. A blurhash is generated for each image; a blurry placeholder remains in place until the image loads. The width, height, and aspect ratio are also recorded.
Videos are converted by default to 1080p H.264 MP4, and also to a 720p version. Initially we used VP9 WebM. Seeing that 4K VP9 videos crashed WebKit on iOS, we switched to H.264. This taught us well the difference between the "most efficient format" and the "format that works on every device".
Alt text: look at the image and write two sentences
The image is resized so it does not exceed 1024 pixels and sent as a JPEG to the Gemini 2.5 Flash‑Lite model on Vertex AI. What we ask of the model is short: to describe the main subject and visual elements of the image in one or two sentences, adding nothing else. Resizing lightens the request and keeps the cost low. Full resolution is not needed to generate alt text.
From description to understanding: vectors
The generated alt text is converted into a 768‑dimensional vector using the bge-base-en-v1.5 model on Cloudflare Workers AI. The vector, together with the alt text, is stored in Postgres with pgvector and searched using an HNSW index for cosine similarity.
We vectorize the description, not the image itself. A single model call does two jobs: the same text is used as alt text on the site and also forms the basis of the search.
Search and the most important rule
When a user searches for "team photo in front of the office", the API converts the query into a vector using the same model and returns the nearest images above a certain similarity threshold.
The most important rule in this pipeline is that both sides use the same model. If the processing service produced vectors with one model and the API with another, nothing would throw an error. The search would silently return meaningless results. Such errors are the most dangerous: everything appears to work but it works incorrectly. Therefore the model name is hard‑coded in both services, and there is a note reminding of this rule.
Privacy: those that never go to the model
Not every file goes through this pipeline. When a business marks a folder as private (e.g., a customer's unpublished photos, a résumé submitted via a form), the files in that folder are not sent to any model and no vectors are created. Alt text is also not generated for visitor‑uploaded files. There is a second safeguard on the search side: the query returns only publicly accessible files.
AI agents can also search
The same search is also exposed to AI clients such as Claude via Marcapony's MCP server. While Claude is preparing a page, it searches for "the shop's showcase photo", finds the image, and inserts it directly into the section. It doesn't need to know the file name.
A known limitation
Alt text is currently generated in English, and the vector model we use is English‑centric. Turkish searches work but are not as accurate as English. Multilingual alt text and search are not yet available; this is the next natural step for the pipeline.