top of page
Abstract Shapes

INSIDE

PUBLICATIONS

The AI Stack 2026: What Every Developer Needs

The AI Stack 2026: What Every Developer Needs
The AI Stack 2026: What Every Developer Needs

UIT emblem

UIT University 365 Institute of Technology

Series AI Engineering | Level Basic (Free)

Duration 15 to 20 minutes | Access Free

IT Engineering, AI and Applied AI, Data Science, Software Development, Digital Transformation


UNOP isochrone

UNOP Sound (University 365 Neuroscience Oriented Pedagogy)

Take five minutes to prepare your brain. Play the isochronous tone track (40Hz gamma frequency) with your eyes closed. Gamma-frequency tones before a learning session raise attention and make the material easier to absorb.

[Audio player: UNOP Pre-Lecture Isochrone (40Hz, 5 minutes)]

In this Lecture


Back to the TOC

The Stack You Actually Ship


You ship a feature that answers questions from your company's documents. It works in the demo. Three weeks later you cannot explain the bill, you cannot tell whether last Tuesday's prompt change made answers better or worse, and one user's runaway loop burned a month of budget in an afternoon.


None of those failures were model failures. They were architecture failures, and they cluster in the layers that sit around the model. The model is one layer of six. This lecture walks the whole stack, layer by layer, with the decision each layer forces and the symptom that appears when you skip it. By the end you will have a stack you can defend, a short list of what to build first, and the four levers that control almost all of your cost.


The seven layers at a glance


USER App and API layer: auth, rate limits, user experience Model gateway: routing, retries, caching, budgets, fallback Orchestration: prompt assembly, tool calls, agent loops Queues and workers: slow and bursty work off the request path Data: system of record, vector store, cache Observability and evaluation: traces, evals, cost dashboards Guardrails: budgets, content checks, tenancy scoping


Each layer has one job. The seams between them are where you place retries, budgets, and checks. A stack that respects the seams survives a model release; a stack that ignores them needs a rewrite every quarter.

Back to the TOC

Layer 1: The Model Gateway


The seven layers of the 2026 AI application stack, with guardrails spanning every layer and the role of each layer
The seven layers of the 2026 AI application stack, with guardrails spanning every layer and the role of each layer

The most expensive mistake in an AI application is coupling. A codebase that calls one vendor's SDK from fifty places cannot switch when that vendor deprecates a model, changes prices, or has a bad week of uptime.


Put a gateway in front. One layer routes requests to the right model, handles authentication and retries, enforces budgets, applies caching, and fails over to an alternate provider.


app code -> gateway.complete(request) -> provider


Swapping models becomes a configuration change rather than a project. Three responsibilities belong here and nowhere else:


  • Routing. Send the hard reasoning to a frontier model and the classification, extraction, and tagging to a small one. Most production traffic does not need a frontier model, and routing it there is the single largest avoidable cost.

  • Fallback. On a provider timeout or error, try the next provider before failing. Every serious provider has bad minutes.

  • Metering. Count tokens and cost per tenant per request. You cannot control a number you do not record at the point of use.


The symptom when you skip it


Your application code contains provider-specific parameters in twenty files, and a price change becomes a two-week task.

Back to the TOC

Layer 2: Orchestration


Orchestration is where the AI logic lives: prompt assembly, tool calls, agent loops, and multi-step workflows. Keep it explicit and testable, and keep it thin.


Workflows before agents, by default


This is the most consequential choice in the layer. A deterministic chain of model calls is predictable, debuggable, and cheap. A free-roaming agent is powerful and introduces variance and cost that you then have to contain.


Start with the simplest orchestration that solves the problem and add autonomy only where it earns its keep. Most production use cases need a workflow, not an agent: retrieve, assemble, call once, validate, return. Reach for an agent loop when the number of steps genuinely cannot be known in advance, such as a task that must search until it finds an answer.


Keep domain logic out


Do not let the orchestration framework grow to own your business logic. Teams that invert this end up unable to test or debug their AI behaviour without invoking the model, which turns every bug into an experiment and every experiment into a bill. The framework assembles context and calls tools. Your application decides what the rules are.


Tool interfaces as a stable boundary


If your agents call tools, define those tools behind a stable interface. A protocol boundary such as the Model Context Protocol lets the same tool serve multiple models and frameworks, and it keeps your integrations from being rewritten when you change the reasoning layer.

Back to the TOC

Layer 3: Queues and Workers


Where a failure shows up when a stack layer is skipped, with the symptom and the incident each omission causes
Where a failure shows up when a stack layer is skipped, with the symptom and the incident each omission causes

AI work is slow and bursty. A generation can take many seconds, traffic spikes, and provider rate limits throttle you. Doing that work inside a web request is a recipe for timeouts and a poor experience.


Push it to a queue:


request -> enqueue job -> 202 Accepted + job id worker pool -> respects provider rate limits, retries, writes result notify -> push the result when it is ready


Three rules make this layer work:


  • Enqueue, do not block. Accept the request, create a job, return a handle immediately.

  • Rate-limit at the worker. Concentrate provider rate-limit handling in the worker pool, so a traffic spike queues gracefully instead of erroring out.

  • Retry with backoff and a dead-letter queue. Transient provider errors are normal. Route persistent failures somewhere visible instead of losing them.


The symptom when you skip it


Your p95 latency is measured in tens of seconds, and a provider slowdown becomes an outage of your own product.

Back to the TOC

Layer 4: The Data Layer


Three distinct stores serve three distinct jobs, and conflating them causes pain that shows up months later.


System of record: a relational database. Users, jobs, results, conversations, audit, billing. This is the source of truth.


Vector store: embeddings for retrieval. Start with the vector extension of the database you already run. For workloads under roughly ten million vectors, that delivers single-digit-millisecond to low-tens-of-milliseconds p95 latency at a small monthly cost, and it keeps your embeddings in the same transaction as your documents and your permission rows. That last property is the real advantage: "only search what this user is allowed to see" becomes a WHERE clause instead of a distributed consistency problem. Move to a dedicated vector database past roughly fifty to one hundred million vectors, or when you need first-class hybrid search and heavy metadata filtering at scale.


Cache: prompt and response caching, rate-limit counters, session state. Caching is a first-class cost-control mechanism, not an optimization you add later. The stable part of your system prompt and your retrieved context are re-sent on every turn; caching them at the provider or in your own store stops you paying full price for the same tokens repeatedly.


Retrieval before fine-tuning


Feeding your data to the model at query time is cheaper to build, cheaper to maintain, and updates the moment your data changes. Fine-tuning changes behaviour, not knowledge, and it is the expensive option. Start with retrieval. Add fine-tuning only when retrieval cannot fix the behaviour you need.

Back to the TOC

Layer 5: Observability and Evaluation


You cannot operate a non-deterministic system you cannot see. Ordinary logging is not enough, because the failure you care about is a quality regression, not a stack trace.


Traces


Log every model call: prompt, model version, output, token counts, latency, and cost, correlated across the steps of one request. Without request-level correlation you can see that a call was slow but not which step caused it.


Evaluation


Build a fixed set of real inputs with expected outcomes, and run it in CI. Fifty representative cases will catch the regression when a provider silently updates a model behind the same name, and nothing else will. Prompt engineering without evaluations is guessing with extra steps.


Cost dashboards


Break cost down by feature and by tenant. Cost in an AI application is an architecture concern: it is driven by how much context you send, how many steps an agent takes, whether you cache, and which model handles which job. Bolting cost control on at the end is painful, because by then the decisions are structural.


Two policies you need before you log anything


Decide your PII and retention policy for prompts and outputs before the first trace lands in storage: prompts routinely contain customer data. And decide what happens when an evaluation fails, or the eval set becomes a dashboard nobody acts on.


The symptom when you skip it


You learn about quality problems from your users, and you have no way to tell whether last week's change helped.

Back to the TOC

Layer 6: Guardrails


Guardrails sit across the layers and enforce the limits you cannot trust application code to remember.


  • Per-user token and spend budgets, enforced server side. A hard daily cap per account. One scripted user can otherwise spend a month of budget in an afternoon.

  • Tenancy scoping. Every retrieval path must filter by the caller's permissions. In multi-tenant products this is a data-layer property, not a prompt instruction.

  • Input and output checks. Validate structured output against a schema before you act on it. A model that returns prose where you expected JSON will fail somewhere downstream, and you want that failure at the boundary.

  • Confirmation before irreversible actions. Anything that spends money, sends a message, or deletes a record goes through a human decision, or at minimum an idempotency key with an audit trail.


The symptom when you skip it


Your first incident is also your first guardrail.

Back to the TOC

The Four Cost Levers


The four cost levers of an AI application: model routing, context size, batching and step count, with the move for each
The four cost levers of an AI application: model routing, context size, batching and step count, with the move for each

Almost all of your spend is controlled by four decisions. Make them deliberately, and measure each one.


Lever

What it controls

The move

**Model routing**

Cost per call

Small model for classification, extraction and tagging; frontier model only for hard reasoning

**Context size**

Tokens per call

Retrieve fewer, better chunks; cache the stable system prompt and retrieved context

**Batching**

Unit price

Anything not user-facing goes through the batch path, which is typically a flat discount

**Step count**

Calls per request

Prefer a workflow to an agent; cap agent iterations explicitly


Three numbers to instrument from day one


Record cost per request, tokens per request broken down by feature, and steps per request. When a bill surprises you, one of those three will explain it, and you will find it in minutes instead of days.


Latency follows the same shape


A response that takes four seconds when it should take one is a churn problem. The fixes are the same levers: caching, routing easy calls to smaller models, and cutting context. Latency is an architecture outcome, not a hosting setting.

Back to the TOC

What to Build First: A Starter and a Scale Stack


Two stacks, and most builds migrate between them piece by piece rather than choosing once.


Layer

Starter stack

Scale stack

Infrastructure

Hosted model API plus a serverless app

Multi-provider routing; self-hosted inference where volume justifies it

Data

Relational database with a vector extension

Dedicated data platform plus a vector database for retrieval scale

Model

One hosted model

Routed multi-model: cheap tier for routine calls, frontier for hard ones, fine-tuned where it pays

Orchestration

Direct API calls through a thin gateway

Graph-based workflows with human checkpoints

Application

Standard web app

Multi-tenant, integrated with existing systems

Observability

Basic logging plus a trace tool

Full evaluation, tracing, drift detection, cost dashboards


The sequence that keeps cost and risk down


  • Start with the use case, not the tools. Define the one job the AI does and what "correct" means for it. The stack follows from that definition.

  • Prove the idea with one hosted model. One API, one prompt, no infrastructure. Confirm the model can do the job before you build around it.

  • Add retrieval when the model needs your data. This is where most products find their real value.

  • Add the evaluation layer before you scale. Wire in tracing and automated checks so you can trust the output as volume grows.

  • Move pieces in-house only when the numbers say so. Revisit self-hosting, a dedicated vector store, or fine-tuning once you have real usage data.


The test for any layer


Ask what happens to your application on the day that component is deprecated or triples in price. If the answer is "we change a configuration file", the layer is built correctly. If the answer is "we rewrite the feature", the layer is coupled, and that coupling is the actual defect.

Back to the TOC

Feynman Summary: Explain It Like You Are 12


Imagine you open a small restaurant. The recipe is not the hard part. What is hard is having a kitchen that works when forty people arrive at once.


You need someone at the door taking orders quickly, so nobody waits in a line (that is the gateway, it decides which cook handles what). You need a kitchen that follows a written recipe every time, rather than a cook who invents something (that is orchestration; a fixed recipe is easier to trust than free improvisation). You need a ticket rail where orders wait their turn, because you cannot cook everything instantly (that is the queue). You need three kinds of storage: the ledger with the real records, the pantry you reach into while cooking, and a shelf of pre-made sauces so you are not remaking the same base forty times (that is the database, the vector store, and the cache). You need someone tasting the food now and then to check it still tastes right (that is evaluation). And you need a rule that no single table can run up an unlimited tab (that is the guardrail).


Notice that none of those jobs is the recipe. The recipe is just one thing in the kitchen. Most AI products fail because of the kitchen, not the recipe.

Back to the TOC

Mindmap: The Complete Picture


Complete mindmap of the 2026 AI application stack: six layers, the decision each forces, the failure symptom when it is skipped, and the four cost levers
Complete mindmap of the 2026 AI application stack: six layers, the decision each forces, the failure symptom when it is skipped, and the four cost levers

The mindmap puts the request path down the centre and branches to each layer, the decision that layer forces, the symptom of skipping it, and the cost levers that cross all of them.



UNOP isochrone

UNOP Sound (University 365 Neuroscience Oriented Pedagogy)

Take five minutes to consolidate your memory. Play the isochronous tone track (10Hz alpha frequency) with your eyes closed. Alpha-frequency tones after a learning session support consolidation, helping move what you just learned from short-term to long-term memory.

[Audio player: UNOP Post-Lecture Isochrone (10Hz, 5 minutes)]

Back to the TOC

Practical Exercise: Map One Request Through the Stack


This exercise takes about 20 minutes on paper, and it is the highest-value planning you can do before writing code.


Step 1: Pick one real feature


Choose a feature you have built or intend to build. Write one sentence describing the single job it does, and one sentence describing what a wrong answer would cost.


Step 2: Trace one request through the seven layers


For each layer, write what it does in your feature, or write "absent" and decide whether the absence is acceptable.


  • App and API: who can call it, and what limits apply per user

  • Gateway: which model, what fallback, where cost is metered

  • Orchestration: how many model calls, and whether the step count is known in advance

  • Queues: which work leaves the request path, and what the user sees while it runs

  • Data: system of record, retrieval source, what is cached

  • Observability: what you log, and the fifty cases you would test on

  • Guardrails: budget cap, tenancy filter, schema validation, what needs confirmation


Step 3: Name the seam


For each boundary between layers, write where the retry, budget, or check lives. A seam with no owner is where the incident will happen.


Step 4: Choose the four cost levers


For your feature, write the model routing rule, the retrieval size, whether anything can be batched, and the cap on steps. Estimate cost per request from those numbers.


What to look for


If you cannot finish Step 2 without inventing components, you have found your build order. If the step count in Layer 2 is unknowable, you have found where an agent is genuinely required, and everywhere else a workflow will do.


Applied AI connection


This is the review you run on any AI feature before it ships, and again after the first month of real traffic. The mapping takes twenty minutes and prevents the rewrite that takes a quarter.

Back to the TOC

Glossary


Term

Definition

**Model gateway**

A single layer that routes model requests, handles auth and retries, meters cost, and fails over between providers.

**Orchestration**

The code that assembles prompts, calls tools, and sequences model calls.

**Workflow**

A deterministic, pre-defined sequence of steps. Predictable and cheaper than an agent loop.

**Agent loop**

A model-driven loop in which the number of steps is decided at run time. Powerful, variable, and costly.

**Queue**

A store of pending jobs that moves slow work off the request path so the API returns immediately.

**Worker**

A process that consumes jobs from a queue, respecting provider rate limits and retrying transient failures.

**Dead-letter queue**

Where jobs go after their retries are exhausted, so failures stay visible instead of disappearing.

**System of record**

The relational database that holds users, jobs, results, and audit as the source of truth.

**Vector store**

A store of embeddings used for similarity search in retrieval.

**pgvector**

A Postgres extension that adds vector columns and similarity search to an existing relational database.

**Prompt cache**

A cache of the stable, repeated part of a prompt so you do not pay full price to resend it each turn.

**Trace**

The recorded path of one request across every model and tool call, with prompt, output, tokens, latency, and cost.

**Evaluation set**

A fixed set of real inputs with expected outcomes, run in CI to detect quality regressions.

**Regression**

A drop in answer quality after a change to a prompt, a model version, or retrieved data.

**Guardrail**

A server-side limit or check that constrains what a request can spend or do.

**Token budget**

A hard cap on tokens or spend per user or per day, enforced outside the model.

**Tenancy scoping**

Filtering every retrieval by the caller's permissions so one customer never sees another's data.

**Idempotency key**

A client-supplied identifier that lets a retried request be recognised instead of executed twice.

**Batch API**

A non-interactive submission path that trades latency for a lower unit price.

**Fine-tuning**

Additional training that changes a model's behaviour. It does not reliably add knowledge; retrieval does.

**p95 latency**

The response time below which 95 percent of requests complete. The number users actually feel.

Back to the TOC

Quiz: TEST YOUR UNDERSTANDING


1. What is the primary reason to put a model gateway in front of your application?


A) It makes the model produce better answers


B) It gives one place for routing, retries, cost metering and provider fallback


C) It removes the need for a vector store


D) It reduces the model's token prices


2. When should you reach for an agent loop rather than a fixed workflow?


A) Whenever you want better quality


B) When the number of steps genuinely cannot be known in advance


C) Whenever the task involves retrieval


D) Always, because agents are more capable


3. Why push long AI work onto a queue instead of running it inside the web request?


A) Queues are cheaper to host


B) Generations take seconds and can spike, so blocking the request causes timeouts and poor experience


C) Queues remove the need for rate limiting


D) The model API rejects requests made from a web server


4. Which statement about retrieval and fine-tuning is correct?


A) Fine-tuning adds knowledge more cheaply than retrieval


B) Retrieval updates the moment your data changes; fine-tuning changes behaviour and is more expensive


C) They are interchangeable


D) Fine-tuning is required before retrieval can work


5. What is the single most valuable evaluation practice in this lecture?


A) Reading the model's release notes each month


B) A fixed set of real inputs with expected outcomes, run in CI on every prompt or model change


C) Asking the model to rate its own answers


D) Comparing vendor leaderboards



Answers: 1-B, 2-B, 3-B, 4-B, 5-B

Back to the TOC

Related Resources


U365 INSIDE Publications



External Resources



Related U365 Lectures (Coming Soon)


  • Multimodal AI: When Models See and Hear, in the AI Frontiers series

  • Model Quantization: Running LLMs on Your Laptop, in the AI Engineering series

Back to the TOC

U.Copilot for This Lecture


Use this prompt with your own AI assistant to review the stack behind a feature you own, and to get a build order out of the result. It follows the UP-Context method: context first, then the task, then the constraints, then the output shape.


CONTEXT I am reviewing the architecture of one AI feature I own. The feature is: [one sentence] The single job it does: [one sentence] What a wrong answer costs: [state the consequence] Expected volume: [requests per day] Current stack: [list what exists today, or write "nothing yet"] Current monthly AI spend: [number, or "unknown"] TASK 1. Trace one request through these seven layers and say for each whether it exists in my feature: app and API, model gateway, orchestration, queues and workers, data, observability and evaluation, guardrails. Mark each as present, partial, or absent. 2. For every layer marked absent or partial, state the failure symptom I should expect, in one sentence. 3. Give me a build order: the three changes with the highest ratio of risk removed to effort spent. 4. Name the four cost levers for my feature with a concrete setting for each: model routing rule, retrieval size, what can be batched, and a cap on steps per request. CONSTRAINTS Use only the information I gave you plus what is in this conversation. Do not invent benchmark numbers or prices. Where a number is needed and missing, say which measurement would produce it. OUTPUT A table with one row per layer: layer, status, evidence from my description, and the symptom if it stays absent. Then the build order as a numbered list of three items, one line each. Then the four cost lever settings as a short list. End with the first three things I should instrument.

Back to the TOC

Next Steps


  • Run the mapping exercise on one feature you own, and write the seam owner for every boundary between layers

  • Add the gateway before you add the second model. Routing without a gateway turns into a rewrite

  • Build the evaluation set now. Fifty real cases with expected outcomes, in CI, before you tune another prompt

  • Set a per-user budget today. It is a few hours of work and it removes your largest uncontrolled risk

  • Read the previous lecture in this series, "RAG vs Fine-Tuning: When to Use Each", if you have not settled your retrieval layer yet


The model is a commodity you rent. Your advantage lives in the layers around it: how well you route, how well you cache, how well you check, and how well you feed it. Build the kitchen, not just the recipe.

Back to the TOC

IMPORTANT NOTICE


This lecture is published by University 365 as part of its INSIDE Publications Hub. The content is free to read for all visitors. Lectures in this series may be part of a structured academic program leading to a Micro-Credential for your Career (MCC). To enroll in an academic program, visit university-365.com/tuition.


This content is for educational purposes. While we strive for accuracy, AI is a fast-moving field. Verify current technical details against primary sources for professional applications.


Copyright University 365, Inc. All rights reserved. This content is protected under University 365's copyright policies. For permissions or inquiries, contact uda@university-365.com.



Published by the Department of Academics, University 365.

Lecture delivered by the University 365 Institute of Technology (UIT).

Sam Utteker, Dean of Technology, UIT

Signed for the academic year 2026.

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
Image by Erik  Lucatero

Become Superhuman

Master AI to stay irreplaceable in every field.

 

 

 

​

​

Apply for Admission Today.
Select Your Initial Access Level.


Become a DISCOVERY, INSIDER, or SUPERHUMAN Fellow.

Image by Milad Fakurian

Master Your Life with a Digital Second Brain

Turn overwhelm into clarity with LIPS + CARE
U365’s unique framework to organize your goals, projects, and knowledge into a superhuman system for success

bottom of page