How Nala is built: architecture, safety, and a layered AI system
After explaining what Nala is and why it makes sense as a product, the technical side deserves its own space. It also helps to remember what the acronym stands for: Natural Adaptiv…
Nala is not designed as a simple call to a model. The interesting part is the architecture around the LLM: a dedicated runtime that prepares context, applies safety, controls tools, executes Genkit, validates output, updates memory and records traceability.
This matters because an AI demo can survive with a large prompt and a direct SDK call. A real product cannot. Once memory, family rules, streaming, parent signals, tools and safety appear, clear boundaries become necessary.

A first AI integration often starts like this:
const response = await openai.chat.completions.create(...)
That is enough to validate an idea, but it does not answer product questions:
Nala separates those concerns into layers. The API adapts HTTP. The runtime governs execution. Genkit orchestrates prompts and the model. Adapters encapsulate persistence and context. Policies control safety, memory and tools.
The main endpoints are:
api/nala/chat.ts
api/nala/stream.ts
Their responsibility is deliberately small: receive the request, validate the public contract, build the internal input and delegate to the runtime.
The normal endpoint ends in:
defaultNalaRuntime.runChatFlow(...)
The streaming endpoint ends in:
defaultNalaRuntime.runChatStreamFlow(...)
Delivery changes, but the execution model remains the same.
The center lives in src/runtime.ts:
createNalaRuntime(deps)
This function creates a runtime with explicit dependencies:
allowLlmToolUse flag;runChatFlow and runChatStreamFlow.This avoids global-singleton coupling. An isolated runtime can use injected adapters and preserve clear guarantees in tests, separate environments or future multi-tenant configurations.
The default runtime uses inMemoryNalaContextAdapter and enables allowLlmToolUse: true for Developer UI and global flows. Custom runtimes default to allowLlmToolUse: false.
A turn is not “prompt → model → response”. It is a pipeline.

The main steps are:
const trace = createTrace('nalaChatFlow', input.metadata?.traceId)
const prepared = await prepareNalaTurn(input, trace, services)
const promptInput = buildChatPromptInput(input, prepared)
const relevantTools = selectRelevantTools(input, prepared, services)
const fallback = fallbackPromptOutput(input, prepared)
const { skip, reason } = shouldSkipChatLlm(prepared)
If the LLM is not skipped, the runtime executes:
const response = await nalaChatPrompt(promptInput, {
tools: relevantTools.actions,
maxTurns: 3,
returnToolRequests: true,
})
And finally:
return finalizeNalaTurn({ input, prepared, promptOutput, trace, services })
The fallback is prepared before the model call. If the LLM fails or safety suggests avoiding it, the system can still return a controlled response.
Genkit is initialized in src/genkit.ts:
export const ai = genkit({
name: 'nala-ai-api',
plugins: [openAI()],
model: `openai/${resolveNalaModelName()}`,
})
In Nala, Genkit provides:
definePrompt;But Genkit does not own all business logic. The healthier architecture is to use Genkit as the AI execution layer while the runtime keeps product decisions.
Contracts live in src/contracts/nala.contracts.ts.
Key schemas include:
NalaChatInputSchema;NalaChatPromptInputSchema;NalaChatPromptOutputSchema;NalaChatOutputSchema;NalaSafetyOutputSchema;NalaIntentOutputSchema;NalaSessionUpdateSchema.This reduces ambiguity. The API does not receive arbitrary input. The prompt does not consume an improvised object. The final response is validated before reaching the client.
In AI systems this matters because model failures can be unusual. Contracts reduce the blast radius.
The right way to view Nala is as a control plane around the LLM.

The model is surrounded by four boundaries:
The LLM answers, but it does not decide by itself what context it sees, what tools it can use or what gets persisted.
Tool calling is one of the most delicate parts of any LLM architecture.
Nala separates two concepts:
createNalaTools(adapters)
creates internal tools closed over injected adapters.
Genkit tools are declared separately with ai.defineTool.

Per-turn selection happens through:
selectRelevantTools(input, prepared, services)
In addition, allowLlmToolUse controls whether actions are actually passed to the model.
This avoids a subtle problem: Genkit tools are global. If they are passed into an isolated runtime carelessly, they may end up using global adapters instead of injected adapters. That is why custom runtimes do not allow LLM tool use by default.
Safety should not be only a line inside the prompt. In Nala it is part of execution.
prepareNalaTurn integrates intent, safety, memory and family rules before the final prompt is built. The safety flow can return:
safetyLevel;allowed;categories;blockedReason;redirectionStrategy;parentSignal.This allows the runtime to decide whether to call the LLM, use fallback or record a parent signal.
For a child/family assistant, that separation is not optional. Privacy, tone and boundaries are part of the product.
Nala does not try to remember everything.
The strategy is to preserve continuity without turning memory into a dangerous container:
Useful memory is usually small. Remembering “she likes dinosaurs” can improve the experience. Storing sensitive personal data should not.
Streaming is implemented as a variant of the same runtime.

Execution calls:
nalaChatPrompt.stream(promptInput, {
tools: relevantTools.actions,
maxTurns: 3,
returnToolRequests: true,
})
Chunks are emitted progressively and, at the end, the complete structured response is awaited, parsed and finalized.
This avoids a second architecture for streaming. Normal and streaming paths share preparation, safety, tool selection, fallback and finalization.
Nala records traces with:
createTrace;addTraceStep;recordToolCall;completeTrace.This shows detected intent, safety level, selected tools, fallback usage and how the turn completed.
In AI systems, observability is not optional. Without traces, debugging bad answers becomes guesswork.
The project includes tests and evaluation datasets for:
That is a good sign. AI applications should not be validated only by trying a few prompts manually. They need repeatable regression cases.
This architecture increases control and testability, but it adds more moving parts:
For a sensitive assistant, that cost is justified. The alternative —one huge prompt with too many responsibilities— scales worse.
Natural next steps would be:
The interesting part of Nala is not only that it uses Genkit.
The interesting part is how it uses Genkit: inside a runtime that controls context, safety, memory, tools, fallback, streaming and final output.
That design is what separates a chatbot demo from a serious foundation for a production AI assistant.