Forum Discussion

Wiliam_Rosa's avatar
Sep 29, 2026

Not Every AI Decision Needs a Chat: Running a Pure Decision Model Straight from SQL

Anyone who has ever pointed a large language model at a bulk classification job, ticket triage, content moderation, queue routing, knows the discomfort: you're paying for a model that can write poetry just to choose between "urgent" and "not urgent". A general reasoning model is great when the question is open-ended, but when the right answer already sits among a handful of known options, it turns into expensive, slow overkill.

A newer class of model attacks exactly that problem: instead of generating free text, it takes a discrete set of options and returns a calibrated decision, complete with a probability per option, at a fraction of the cost and latency of a chat model. And the detail that matters to anyone working with data: you can query this kind of model directly from SQL, against tables already governed by Unity Catalog, without writing a single line of agent orchestration code.

What changes when the model only needs to decide, not converse

A decision model isn't a smaller version of an LLM, it's a different category built for a different purpose. It doesn't hold conversation state, doesn't generate a natural-language explanation by default, and isn't trying to be useful for just any question. It solves a narrow problem: given a fixed set of possible labels, which one is most likely for this input, and with how much confidence.

That matters because most real-world classification work in production, moderation, routing, triage, risk scoring, is already fundamentally this kind of closed decision. Running a general reasoning model for that kind of task is like hiring a senior consultant to stamp a form: it works, but the cost per decision stops making sense at scale.

ai_query: the bridge between the model and data that already lives in Unity Catalog

Azure Databricks exposes "ai_query", a general-purpose AI Function that queries any supported model directly from SQL or Python, no separate inference pipeline required. What sets it apart from task-specific AI Functions ("ai_classify", "ai_summarize", and similar) is control: "ai_query" lets you choose the model, the prompt, and the parameters manually, exactly what you need to point at a custom decision model hosted on your own Model Serving endpoint.

Two requirements worth checking before you try this: "ai_query" doesn't run on Classic SQL warehouses, it needs Serverless; and the minimum Databricks Runtime is 15.4 LTS, with 18.2 or above recommended for the best performance.

Hands-on: querying a custom decision model

A custom decision model, in the Jev family, typically doesn't expose a standard chat-completions interface, it's served as a regular Model Serving endpoint, and the "ai_query" call uses the custom-model syntax, with explicit "endpoint", "request", and "returnType":

SELECT
  ticket_id,
  description,
  ai_query(
    endpoint => "support-triage-decision",
    request => named_struct(
      "text", description,
      "channel", source_channel,
      "customer_plan", plan
    ),
    returnType => "STRUCT<category: STRING, urgency: STRING, confidence: DOUBLE>"
  ) AS decision
FROM support.open_tickets
WHERE status = 'new';

The structured return ("category", "urgency", "confidence") is already shaped to become a table column, no free-text parsing or regex needed to pull a field out of a chat response. It's the same pattern the official documentation uses for a traditional ML model (their example is a spam classifier), just swapping the endpoint for the decision model.

My take: the win here isn't just cost per call, it's architectural. Putting the decision inside SQL itself means the result is born as a column in a governed table, right where Unity Catalog's lineage, access permission, and change history already apply. Compared to a separate agent pipeline that writes back to the lakehouse afterward, that's one less failure surface.

Three ways to host the model, three levels of operational effort

"ai_query" accepts three categories of model, and the choice between them changes how much infrastructure someone ends up maintaining:

Platform-hosted model ("system.ai.*"), like the Claude, Llama, and Qwen families already offered as a service. Nothing to provision, it scales on its own, and it's the recommended option for anyone who just needs batch inference without customizing model weights.

Provisioned-throughput model, for anyone who has already fine-tuned a foundation model or needs reserved capacity. Here the Model Serving endpoint has to be created and sized manually, but "ai_query" still handles the parallelization and retries of the batch query, it doesn't use the endpoint's own capacity for that.

Custom or external model, the category a decision model like Jev falls into: a traditional ML model, trained outside the foundation-model ecosystem, or hosted outside Azure Databricks via an external model endpoint. It's the option with the most control, and also the one that demands the most infrastructure decisions on your own, creating and maintaining the serving endpoint is on whoever operates it.

For anyone deciding where to invest platform effort, that list already works as a priority order: start with what the platform hosts for free, and only move to the custom option once the specific task justifies the extra work of operating the endpoint.

Unity Gateway governance, with a real caveat

A call to "ai_query" against a hosted model ("system.ai.*") is automatically routed through Unity Gateway, but only a subset of gateway functionality applies to that route: usage tracking and budget integration (including hard spend caps) work normally, but guardrails via service policy, inference tables, tracing tables, rate limits, and fallback between providers don't apply to a call routed through "ai_query". It's easy to assume that's "already included" and find out otherwise in production.

What this doesn't solve on its own

"ai_query" solves how the call gets there, not the quality of the decision. A custom decision model still depends on training curation and evaluation against ground truth, the documentation itself recommends using Agent Evaluation to measure batch accuracy before trusting the result in production. And unlike the platform's natively hosted "system.ai" models, a fine-tuned or fully custom model still requires provisioning and maintaining the Model Serving endpoint yourself, without the automatic autoscaling that natively hosted models get for free.

Bottom line

Not every AI problem is a conversation problem. When the output is a decision within a closed set of options, a decision model queried via "ai_query" tends to cost less, respond faster, and integrate better with data that already lives in Unity Catalog than standing up one more conversational agent to do the same triage. In practice: the question worth asking before writing the next classification prompt is simple, is this a decision among known options, or is it genuinely an open-ended question? The answer completely changes which kind of model is worth paying for.

References

No RepliesBe the first to reply