Home Articles Resume Nala Project
ES EN

A small model as the first stop

Illustration: many lines converging through an orange filter before reaching a dark block, with the text “Decidir antes de generar. Menos contexto. Menos espera.”

With Laya and Jev, something happened to me that hadn't happened with a model in a long time. I thought: this is exactly what I was trying to achieve.

In our Gateway, we had set up Gemma 4 E2B as the first stop. A small model with a very narrow task: look at the request and help us decide what to do with it before calling a bigger model. Whether we needed to call a bigger one at all, or which decisions to make.

There was quite a lot of work behind it. We had to define the options, explain clearly when each one applied and work out what to do with the queries that didn't fit.

What we were after was fewer calls to the generative LLM. It should step in when something needs writing, explaining or reasoning that's worth it. To pick a tool or decide where to look up a piece of data, I found it hard to justify making it read every instruction and the whole tool catalog each time.

For an SME this is vital. The budget is limited, so is the team, and everything you add has to be maintained later. And personally, what worries me most is the wait: you can optimize the price per token a lot, but if you chain several calls and the user is still waiting, you still have a problem.

That's why the low latency of this kind of model interests me so much. If the first decision is fast and accurate enough, you can save work further down the line: retrieve only the information you need, offer the LLM fewer tools, and avoid context that adds nothing and only slows responses down and increases the input tokens to process. You can also add checks that help with guardrails, while keeping permissions and authorization in the Gateway's code.

But what I'm most eager to explore is fine-tuning, which is already available in Laya.

I want to be able to train it on our own cases, reviewed by the people who know the work. What we mean by an urgent request. When it should go to another team. What information we need before moving on. Which exceptions need someone to stop and take a look. And since it's a small model, it's genuinely feasible to do this on an SME's infrastructure.

Those criteria are usually spread across procedures, the experience of a few people and technical documentation scattered around the company's ecosystem. Being able to capture them as examples and check whether a small model learns to apply them seems especially valuable to me for an SME. Of course, first you have to agree on what they are: if we resolve two equivalent cases differently, training isn't going to fix that by magic.

Beyond assistants, I see potential in applications where a decision has to arrive in milliseconds. There you'll need to measure the full round trip, because the network and the rest of the services count too.

There's still testing to do, to see where it works well and where it fails. A model returning a valid option doesn't mean it picked the right one.

Even so, it had been a while since a line of work fit so well with what I needed. If we manage to adapt these decisions to our own criteria and resolve them with very low latency, I think we'll be able to do a lot more with the resources we have. For many SMEs, that could change things a lot.

#Artificial Intelligence#SMEs#Laya#Jev#Fine-tuning#AI Agents#LLM#Software Architecture#Automation#Applied AI

Alex Sanz

I build products and systems where architecture, business needs, and AI become real, reliable, and maintainable capabilities.

Related articles

View all
014 min

ADRs: memory for humans and context for agents

A few years ago, during my time at Santander Bank, we started using Architecture Decision Records (ADRs) to keep track of certain architectural decisions: what had been decided, wh…

032 min

AI, responsibility, and the model we are building

Artificial intelligence is moving incredibly fast and I think that, as a society and as companies, we need to talk about it with more calm, more responsibility, and less euphoria. …