Building a Customer Support Agent with Jev: From a Chatbot to a Decision-Driven AI System

Building a Customer Support Agent with Jev: From a Chatbot to a Decision-Driven AI System

Rudra Prasad Bhuyan

While exploring Jev from TypeSafe AI, I wanted to understand where a decision model like this actually fits inside a real AI application.

A customer-support agent is a useful example because a support message rarely contains only one signal.

A customer might say:

A system looking at this message may need to determine several things at once:

  • What is the main intent?
  • Is the customer asking for a refund?
  • How frustrated are they?
  • How urgent is the problem?
  • Is there a risk of churn?
  • Should this be handled automatically?
  • Should a billing agent, retention agent, or human handle it?

A traditional LLM agent can reason about all of these questions. But the interesting idea behind Jev is to separate decision-making from language generation.

Jev is TypeSafe’s System One model. Instead of primarily producing prose, it evaluates typed questions against some state and returns structured values that application code can directly use. Its three main primitives are Choice, Score, and Noul. (TypeSafe AI)

That leads to the architecture in the diagram.

1. Customer Input: Where Everything Starts

The first component is simple:

Customer → Input channel → Backend

The message might arrive through: a chat widget, email, a mobile application, WhatsApp, support portal, or even a voice system after speech-to-text.

At this point, we have mostly unstructured text.

For example:

“I was charged twice for my subscription and haven’t received a response. I’m really frustrated and might cancel if this isn’t fixed soon.”

An ordinary chatbot could immediately send this message to an LLM and ask:

That can work, but now one model is being asked to perform several jobs simultaneously: classify the problem, judge severity, infer customer sentiment, decide whether tools should be called, decide whether escalation is necessary, and finally generate the response.

For a simple demo, that may be enough.

For a production system, it can be useful to separate those responsibilities.

2. Build Context: Give the Decision Layer the Right State

The customer message alone usually isn’t enough.

Before making a decision, the backend can construct a richer state.

Imagine the same message with additional information:

Customer:
    account_age: 4 years
    plan: enterprise
    lifetime_value: high
Subscription:
    status: active
    renewal_date: Sep 20
Payments:
    Sep 20: ₹2,499
    Sep 20: ₹2,499
Recent tickets:
    billing ticket opened 2 days ago
    no response yet
Message:
    "I was charged twice..."

Now the AI is no longer judging one sentence in isolation.

It has the state of the situation.

This part remains normal application engineering. Your backend might collect the information from PostgreSQL, Stripe, CRM systems, ticket databases, internal APIs, or other services.

Jev does not replace those systems.

It evaluates the state you give it.

3. The Jev Decision Layer

This is the central part of the architecture.

Instead of asking:

“What should we do with this customer?”

we can break that large question into smaller, atomic decisions.

For example:

What is the primary intent?
How angry is the customer?
How urgent is this?
Is the customer requesting a refund?
Is there a meaningful churn risk?
Is this a high-value customer?

TypeSafe specifically recommends this kind of decomposition: ask focused questions independently and combine the answers using application logic when the final decision depends on several factors. (TypeSafe AI)

Jev provides three primitives for doing this.

4. Choice: Pick One Option

The first primitive is Choice.

Use Choice when one answer needs to be selected from a known set.

For our support system:

So the most likely answer is: billing_issue

This distinction becomes important later.

Instead of:

intent = "billing"

we effectively know:

intent = "billing"
confidence = ...
probability_distribution = {...}

Now the backend can decide how much it trusts that classification.

5. Score: Measure Something on a Spectrum

Some questions don’t have categorical answers.

Consider frustration.

We don’t necessarily want:

angry = true

We may instead want something more like:

How frustrated is this customer?

with levels representing something like:

0 → calm
1 → mildly frustrated
2 → frustrated
3 → very frustrated
4 → extremely frustrated

That is where Score fits.

A Score judges something against ordered descriptive levels. Jev can return a position on that scale, along with probabilities for the levels and a confidence value. (TypeSafe AI)

We might evaluate:

anger
urgency
problem severity
churn risk
customer satisfaction

separately.

That is much more useful than hiding everything inside one vague prompt such as:

“How important is this ticket?”

6. Noul: Ask a Yes/No Question

The third primitive is Noul.

A Noul represents the probability that a yes/no statement is true.

For example:

"Is the customer requesting a refund?"

The result could be:

0.92

A value near 1 means a strong yes, a value near 0 means a strong no, and values around the middle indicate ambiguity. (TypeSafe AI)

For customer support, Nouls could ask:

Is the customer requesting a refund?
Is the customer threatening to cancel?
Is human escalation requested?
Is this a repeat complaint?
Is the customer a high-value customer?
Does the message contain sensitive information?

Then our normal Python or TypeScript code decides what those values mean operationally.

For example:

if refund_requested > 0.90:
    require_refund_workflow()

The threshold belongs to the application.

That is an important architectural idea.

The model provides the judgment. The software controls the policy.

7. One State, Multiple Decisions

Now the architecture becomes more interesting.

For the same customer message, we might ask Jev all of these questions:

Choice:
    What is the intent?
Noul:
    Is a refund requested?
Noul:
    Is cancellation being considered?
Score:
    How frustrated is the customer?
Score:
    How urgent is the problem?
Noul:
    Is this a high-value customer?

These questions can be evaluated together. TypeSafe’s documentation describes questions in one request as being evaluated independently and in parallel against the same state. (TypeSafe AI)

So instead of building a long sequential chain:

and then use the results in code.

This leads directly to the patterns shown in the second half of the diagram.

8. Pattern One: Intent Routing

The first pattern is Intent Routing.

Suppose Choice returns:

billing_issue       0.82
technical_issue     0.08
account_management  0.05
product_question    0.03
other               0.02

The application can now route the ticket:

billing_issue
      ↓
Billing Support

Another ticket might go:

technical_issue
      ↓
Technical Support Agent

and another:

account_management
      ↓
Account Service

This is also where the difference from a generic LLM architecture becomes clearer.

A common architecture is:

Customer
   ↓
Large LLM
   ↓
"Figure everything out"
   ↓
Tools

With a decision layer, we can instead build:

Customer
   ↓
Jev
   ↓
Determine intent
   ↓
Backend router
   ├── Billing workflow
   ├── Technical workflow
   ├── Specialist LLM
   ├── Database lookup
   └── Human

Not every ticket necessarily needs an expensive generative model.

A password-status request might only require deterministic code and a database lookup.

A complicated billing dispute might require a specialized LLM with tools.

A highly ambiguous complaint might go directly to a person.

9. Pattern Two: Confidence-Gated Routing

Classification alone isn’t enough.

Suppose the model says:

intent = billing_issue

The next question should be:

How certain are we?

This is where confidence-gated routing enters.

And thresholds do not have to be universal.

For example:

FAQ classification
threshold = 0.60

billing routing
threshold = 0.75

automatic refund
threshold = 0.95

The consequences of being wrong are different.

That policy remains visible in application code rather than being buried inside a long system prompt.

10. Pattern Three: Speculative Fan-Out

Now consider our example again:

“I was charged twice for my subscription and haven’t received a response. I’m really frustrated and might cancel if this isn’t fixed soon.”

There may not actually be only one relevant intent.

We could find:

billing_issue        0.60
refund_request       0.25
subscription_cancel  0.10
payment_method       0.05

This ticket potentially involves:

Billing
Refund
Retention

Instead of waiting for one classification and then making additional AI calls, speculative fan-out asks useful downstream questions upfront.

TypeSafe’s pattern describes sending multiple questions together — even ones that may ultimately prove irrelevant — and then letting normal code ignore the answers it does not need. (TypeSafe AI)

Conceptually:

Suppose the customer never requested cancellation.

The retention-related result can simply be ignored.

But if cancellation risk turns out to matter, the result is already available.

This becomes especially useful in workflows with many possible branches.

11. Pattern Four: Composite Scoring

Sometimes no single signal should decide what happens.

Imagine we have:

anger              = 0.91
urgency            = 0.45
churn risk         = 0.78
high-value customer = 0.85

We might define an escalation score ourselves:

priority =
    0.4 × anger
  + 0.3 × urgency
  + 0.2 × churn_risk
  + 0.1 × customer_value

Giving approximately:

priority = 0.74

Now the business can decide:

0.00–0.40 → normal queue
0.40–0.70 → priority queue
0.70–1.00 → urgent escalation

The interesting part is that Jev isn’t deciding the business formula.

Your application does.

TypeSafe calls this pattern Composite Scoring: evaluate independent dimensions and combine them using weights controlled by code. (TypeSafe AI)

If the business later decides churn matters more, we don’t necessarily need to rewrite one giant prompt.

We can change:

priority = (
    0.30 * anger
    + 0.20 * urgency
    + 0.40 * churn_risk
    + 0.10 * customer_value
)

That makes the final decision logic much easier to inspect.

12. Take Action: This Is Where the LLM Can Return

Jev does not mean removing LLMs from the architecture.

It can actually make their role more specific.

Imagine Jev determines:

intent = billing_issue
confidence = high
refund_requested = high
frustration = high
cancellation_risk = high

Now we know which generative system should be called.

For example:

Jev
 ↓
Billing specialist LLM
 ↓
Tools
 ├── get_invoice()
 ├── inspect_payment()
 ├── issue_refund()
 └── update_ticket()

The LLM can concentrate on what generative models are good at:

  • understanding detailed instructions,
  • working with tool results,
  • reasoning through complicated cases,
  • generating a natural customer response,
  • summarizing the investigation.

Jev handles the smaller structured judgments that determine which path should run.

So this isn’t really: Jev vs LLM

It is more useful to think of it as:

  • Jev → decide
  • Code → control
  • LLM → reason/generate when required
  • Tools → perform actions

13. What a Normal LLM Architecture Might Look Like

Without this separation, we might build:

Modern LLM systems can absolutely produce structured outputs, so this architecture is possible.

The difference is architectural.

Jev is specifically built around structured decisions: Choice, Score, and Noul return values designed to be consumed by software rather than prose designed primarily for a human reader. (TypeSafe AI)

That allows us to move from one large fuzzy instruction toward several explicit decisions.

For example:

Instead of:
"Should this ticket be escalated?"

we have:
anger = ...
urgency = ...
churn_risk = ...
customer_value = ...
intent = ...
confidence = ...

and then:
if priority > 0.70:
    escalate()

That separation is what I find interesting.

14. Learn and Monitor

The final component in the diagram is:

Take Action
     ↓
Learn & Monitor
     ↓
feedback into the system

This shouldn’t be interpreted as Jev automatically learning from every support ticket.

The application should capture what happened after each decision.

For example:

Predicted intent: billing
Actual team: billing
Predicted refund request: 0.92

Human decision: refund requested
Predicted high priority: 0.74
Actual escalation: yes

Customer outcome:
issue resolved
subscription retained

Over time, we can evaluate:

  • which classifications are wrong,
  • which confidence thresholds are too aggressive,
  • which questions are ambiguous,
  • which routes frequently need human correction,
  • whether our composite weights make sense,
  • and where automation should be reduced or increased.

Then we improve the system.

or we rewrite one Choice criterion because billing and subscription-management requests are being confused.

This closes the loop.

Putting the Whole Architecture Together

The complete system becomes:

That is the main idea behind the diagram.

The customer still experiences a normal support agent.

But behind that simple interface, the system has separated understanding, judgment, policy, generation, and execution.

The Main Idea I Took Away From Jev

What interested me most about Jev wasn’t simply getting another model response.

Jev does not need to replace the LLM.

It can become a decision layer around the LLM, helping software determine when an LLM is needed, which LLM or workflow should receive the request, which tools should be exposed, when automation is safe, and when a human should take over.

For a customer-support agent, that turns the architecture from simply “chat with an LLM” into something closer to a controlled decision system.