A small model as the first stop
With Laya and Jev, something happened to me that hadn't happened with a model in a long time. I thought: this is exactly what I was trying to achieve. In our Gateway, we had set up…
I have been experimenting with a multimodal RAG pipeline built on top of images.
What interested me most was not just analyzing the image itself, but turning that output into retrievable context for an LLM.
The vision part is obviously not new. Image analysis, attribute extraction, and content classification have been around for many years. What I wanted to test was something else: how to use that visual information to enrich a RAG system and give a model better context.
That is where the experiment has been focused: take an image, analyze it locally with a vision model, convert that output into something structured, and then prepare it so it can become part of a vector index. In the end, what I want is not just for the model to describe an image, but for that description to become retrievable context.
I have been testing several local models with Ollama and, using the same prompt and the same images, Gemma 4 e4b has given me surprisingly good results, honestly much better than Qwen2.5 VL.
Not because it does anything spectacular, but because of something much more practical: it returns fairly clean JSON, preserves structure well, and produces descriptions that are consistent enough to keep building on.
Most of the time goes into prompt tuning, trying variations, cleaning responses, normalizing fields, preventing the model from over-claiming things it should not assert, and deciding when an output is reliable enough or when it should be flagged for review. That part is less flashy, but it is also what makes the experiment start to look usable.
I especially like the separation of responsibilities. The database remains the source of truth for structured data. Vision adds a semantic layer on top of what can be observed in the image. Embeddings make that information easier to retrieve flexibly. And the LLM receives better context to answer.
That is where I think the value is: not in replacing the data that already exists, but in complementing it.
With this approach, you can start thinking about more natural search over image collections, not only through names, codes, or manual tags, but also through visual features, tone, texture, pattern, finish, or semantic similarity.
Next step: test it with more samples, measure retrieval quality, and combine this textual layer with actual visual search.