Back to blog

Laya: How to Use Typed AI Decisions in Python in 2026

Hello HaWkers, most AI automations do not need to write a paragraph. They need to choose a queue, estimate how urgent a request is, or answer whether a particular risk is present. The open source project Laya has gained attention on GitHub by focusing on exactly this kind of decision. When checked on September 27, 2026, the GitHub API showed more than 26,000 stars for the repository; that number changes continuously.

Suppose a customer sends you a message and your system needs to decide what happens next. Should it call a generative model, parse whatever the model says, and hope the response follows the required format? In this guide, we will build a small triage example with Laya, examine what its output actually means, and define the points where a person should still review the result. The aim is to make the boundary between a model prediction and a business action explicit.

What is a typed decision, and why does it matter?

A typed decision has a known set of possible outputs before inference starts. You might request a choice among billing, technical, and other; an urgency score; or the probability that a statement is true. That structure helps when the next component expects a defined field. It does not have to interpret a free form sentence all over again, guess which part contains the answer, or recover from unexpected prose.

Laya provides three question types in its API: choice selects an option, score evaluates options on a scale, and noul estimates the probability of “yes.” The unusual name noul appears literally in the official examples. Keep that spelling in the API. Replacing it with an intuitive name such as yes_no would mean discovering an integration error only after writing the rest of the workflow.

The project describes its engine as non autoregressive: it evaluates questions without generating a response word by word. This removes some of the work involved in turning free form text back into structured data. A structured response, however, does not guarantee a correct decision. Ambiguous criteria, weak examples, and inputs outside the intended domain can all produce a poor prediction that happens to fit a valid JSON shape.

The project documentation reports 33 ms for a single question on a T4 GPU and 7.2 ms per question in a batch. These are measurements reported by the authors under particular conditions. They are not latency promises for your laptop, server, or real production traffic. Before choosing a tool for speed, measure model loading, input length, hardware, batching, and request queueing in the environment where you will use it.

How Router chooses a model

The interface recommended in the README is Router. It identifies the input language and writing system, then routes the request to an appropriate checkpoint. The repository describes three checkpoints: one for English, one multilingual, and one tuned for typed decision workflows. An initial download needs access to Hugging Face. Subsequent runs may reuse the environment's local cache rather than fetching the same files again.

This detail affects production architecture. A plain Router() loads models as needed. With Router(preload=True), the checkpoints load at startup instead, trading startup time and memory use for less waiting on initial requests. If the application receives messages in several languages, a small cache can lead to repeated model switching. For an API with steady traffic, choose a loading policy deliberately and watch memory consumption before increasing the number of replicas.

Automatic routing is useful, but it does not replace tests by language. A message mostly written in English with one short Portuguese phrase may take a different path from a document written entirely in Portuguese. The documentation also lets you request model="multilingual" explicitly. Use that option when your tests show that automatic selection does not fit the actual messages your application receives. The best checkpoint is a measured choice, not an assumption about a language label.

There is a difference between routing and inference as well. The command line tool, without --predict, can show the routing decision without downloading a checkpoint. That makes it useful for inspecting language selection quickly. It does not prove the classification itself is accurate. To assess accuracy, you need labeled examples and a complete model run that you can compare with the expected answers.

Installation and a first check

According to the README consulted for this article, the package requires Python 3.10 or newer and is available on PyPI. Create an isolated environment so that the package and its dependencies do not change other projects. The commands below follow the documentation for macOS and Linux; on Windows, adapt how you activate or call the environment.

# Create an isolated environment in the project directory.
python3 -m venv .venv

# Install the published PyPI package inside that environment.
.venv/bin/python -m pip install laya

# Confirm the installation without loading a checkpoint.
.venv/bin/python -I -c "import laya; print(laya.__version__)"

Printing the version checks the installation; it does not measure model performance. On the first inference, Router may need to download large files. Plan for that warmup during deployment, especially if your service starts instances on demand. If the runtime cannot reach the Hub, arrange a cache or a checkpoint distribution process before placing the endpoint behind a load balancer. Otherwise, the first real customer request can become an accidental setup step.

Practical example: triaging a request

Now let us turn a message into fields that a support system can use. The team names in this example are local choices. Adapt criteria to the real responsibilities and language of your organization. Each question describes what should be evaluated instead of merely supplying a loose list of labels. That distinction matters because a label may mean something different to the model than it does to the people operating your queue.

from laya import Router

# The first use may download the required checkpoint.
router = Router()
mensagem = (
    "Minha assinatura foi cobrada duas vezes neste mês. "
    "Preciso do estorno hoje ou vou cancelar."
)
perguntas = {
    "equipe": {
        "type": "choice",
        "instructions": "Qual equipe deve receber a solicitação?",
        "criteria": {
            "financeiro": "cobranças, pagamentos e reembolsos",
            "suporte": "falhas, indisponibilidade e erros técnicos",
            "outros": "pedidos que não pertencem às outras equipes",
        },
    },
    "urgencia": {
        "type": "score",
        "instructions": "Quão urgente é a resposta?",
        "criteria": ["pode esperar", "em breve", "bloqueante"],
    },
    "risco_cancelamento": {
        "type": "noul",
        "instructions": "A pessoa ameaça cancelar o serviço?",
    },
}

resultado = router.predict(mensagem, perguntas)
print(resultado["answers"]["equipe"]["choice"])
print(resultado["answers"]["risco_cancelamento"]["noul"])
print(resultado["routing"]["model"])

The customer message says that a subscription was charged twice this month and asks for a refund today to avoid cancellation. The example uses Portuguese field names and criteria to match that input; the choice, score, and noul keys are the API terms that must remain as written. For a real English language workflow, write and test your own English criteria with representative English messages rather than assuming a literal translation gives identical results.

The access pattern at the end follows the official quickstart. The choice field contains the selected option, while noul represents a probability. It is not permission to take an irreversible action. The routing field tells you which model handled the request. For urgency, inspect the complete answer and its distribution before treating the first displayed value as a business priority. A score is evidence for your policy to interpret, not the policy itself.

This separation between prediction and action is fundamental. The model can suggest the finance queue, but your application decides whether to create a ticket, request documents, or ask for a review. It should not infer a refund policy from one sentence written by a customer. Record the input, the model version, the criteria used, and the result so you can investigate mistakes and compare future changes. Handle those records according to your privacy requirements, because customer messages may contain personal data.

How to handle confidence and human review

A probability can look like a ready made decision, but it is useful only when calibrated for your domain. A prediction of 0.8 should correspond roughly to the event's observed frequency among comparable examples. You have to measure that relationship on labeled data. It does not follow automatically from the API returning a number between zero and one. Laya's documentation mentions calibration adjustment and provides fine tuning material, but validating your workflow remains your responsibility.

Start with a simple policy: automate only low impact decisions and send uncertain cases to a person. Do not copy a universal threshold from another product. Choose a threshold after comparing false positives, false negatives, and operational cost on your own validation set. A tool that routes tickets can tolerate a mistake that would be unacceptable for blocking an account or denying access. Reviewers also need enough context to correct a suggestion without having to reconstruct why it was made.

# Use the same response produced in the preceding example.
resposta = resultado["answers"]["risco_cancelamento"]
probabilidade = resposta["noul"]

# This threshold is illustrative only; calibrate it with real data.
if probabilidade >= 0.80:
    encaminhamento = "revisao_humana_prioritaria"
else:
    encaminhamento = "fluxo_normal"

print({"encaminhamento": encaminhamento, "probabilidade": probabilidade})

Even here, the “normal flow” should not mean ignoring the customer. It can mean standard support with a defined owner and response deadline. The code demonstrates how to keep a model estimate separate from routing logic; it does not establish that 0.80 is the right cutoff. You need to see how many real cases fall on either side, what kinds of errors occur, and how the support team handles them.

Also review cases where messages are short, sarcastic, contain several requests, or arrive in a language that appears rarely in your data. These slices often reveal problems that an overall accuracy average conceals. If urgent messages are systematically written in a particular style, an apparently good global metric can still miss the people who most need help. Inspect those failure patterns before you let a predicted value drive an operational decision.

Multilingual tests without assuming equivalence

The project claims support for more than 100 languages in its multilingual checkpoint. That coverage is relevant to a blog and to products with customers speaking Portuguese, English, Spanish, and French. “Supported,” though, does not mean equally accurate in every language. The words used for cancellation, billing, and urgency differ by culture and industry. Assemble labeled examples from each market before automating decisions across them.

The README shows that the same Router can receive a Spanish sentence and route the request to the multilingual model. The test below reuses the team question from the previous example. Its routing output lets you check which checkpoint was selected, while the choice shows what the model predicted. To compare checkpoints fairly, select one explicitly for each run and hold the test set constant.

# Reuse the router and questions defined earlier.
entradas = [
    "Me cobraron dos veces este mes; necesito un reembolso.",
    "J'ai été facturé deux fois ce mois-ci ; je demande un remboursement.",
    "Fui cobrado duas vezes neste mês; preciso de reembolso.",
]

for texto in entradas:
    resposta = router.predict(texto, {"equipe": perguntas["equipe"]})
    print(texto, resposta["routing"]["model"])
    print(resposta["answers"]["equipe"]["choice"])

These three sentences all request a refund after a duplicate charge, but they are only a demonstration of the API path. For a serious evaluation, do more than translate the same examples. Include messages originally written in each language, abbreviations, typing mistakes, and sentences that mix languages. Separate training and testing by customer or time period to avoid leakage from near duplicate examples. Track the share of cases sent to human review by language, too: a model may appear accurate simply because it abstains more often for one group.

How to measure whether it is worthwhile in production

The repository presents its own benchmark for typed decisions, with 2,000 decisions across four workflows. In that benchmark, the tuned checkpoint reached 0.766 accuracy versus 0.362 for the base English model on the same tasks, according to the authors. That illustrates what adapting a model to a problem domain can achieve. It does not establish that every company will see the same improvement. Class distribution, label quality, and the way the criteria are written all have a large effect on the result.

Design your test before comparing Laya, rules, and a generative model. Choose measurements that reflect the business outcome: accuracy by class, a confusion matrix, calibration, high percentile latency, operating cost, and the fraction of cases handed to people. In support operations, measure how long the team takes to correct a wrongly routed request. A fast classifier that creates extra work can end up costing more than a slower approach that makes fewer harmful errors.

Use a simple baseline. Some customer messages contain words and patterns that deterministic rules can handle well; others need context that a classifier does not have. Compare alternatives on the same test set with criteria held fixed. If a prompt or checkpoint change improves the average but makes critical cases worse, do not promote that change automatically. Keep a compact regression set with difficult examples and anonymized real incidents, and run it when you change the decision flow.

If your application already uses generative models, the architecture does not have to choose just one method. A typed stage can select a queue or prioritize review, while a generative stage drafts a reply for an agent to approve. The useful boundary is clear responsibility: classification, text generation, and the final business decision are different operations. For guidance on choosing a generative model by cost and usage profile, see our guide to choosing between GPT-6 Sol and Luna.

Outlook: less free form text, more auditable decisions

Smaller, specialized models make sense when an application repeatedly has to answer the same questions. Laya gives you a concrete interface for testing that idea: a question schema, language routing, structured answers, and checkpoints you can run in your own environment. Attention on GitHub tells us that developers are interested. It does not establish operational maturity or prove quality for your particular use case.

Begin with a reversible workflow, such as suggesting the queue for a ticket. Define clear criteria, collect real examples with consent, and remove unnecessary personal information. Measure mistakes by language and request type. Only then decide whether the automation saves time without hiding problems. Where money, access, or safety is involved, preserve a way to appeal a decision and obtain human oversight.

The code in this article is a starting point for experimentation, not a complete production system. Official documentation can change, and numbers published by the project should be reproduced on your own hardware before they become targets. The more durable architectural lesson is to describe the choice explicitly when you need a choice, measure uncertainty, and keep the action under the control of your software and the people responsible for it.

Let's go! 🦅

📚 Want to Keep Up With What Is Coming?

This article covered typed decisions with Laya, but the ecosystem changes every week, and not every development becomes an article here.

On X, I share what I am testing, the work behind my projects, and new tools and ideas before they turn into blog posts.

Follow Me There

👉 Follow @jeffbruchado on X

💡 Daily content about development, careers, and the tools I actually use

Comments (0)

This article has no comments yet 😢. Be the first! 🚀🦅

Add comments