---
title: Laya is now available through the LayerCloud gateway
date: 2026-09-30T03:01:16Z
canonical: https://ai.layercloud.org/blog/laya-decision-service-now-available
author: LayerCloud
tags: release, laya, models
---

# Laya is now available through the LayerCloud gateway

> A local decision service that answers typed questions with typed answers. What it does, how to call it, the three question types, how companies use it, and where its limits are.

We have added a **local decision service** to the gateway. It answers one kind of question very well, and it is worth being clear about which kind.

## What it does

Laya is a *decision router*. You give it a piece of text - a support message, a form field, a complaint - and a set of **typed questions** about that text. It returns an answer for each question, with a confidence value.

It is not a chat model. It does not hold a conversation, and it will not write your email. Asking it an open question and expecting prose will disappoint you.

## Two ways to call it

**The native endpoint**, which is Laya's own shape:

```bash
curl https://ai.layercloud.ir/v1/systemone \
  -H "Authorization: Bearer $LAYERCLOUD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "state": "I was charged twice this month and nobody has replied",
    "questions": {
      "department": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": ["billing", "technical", "other"]
      }
    }
  }'
```

**The OpenAI-compatible endpoint**, if your code already speaks it:

```bash
curl https://ai.layercloud.ir/v1/chat/completions \
  -H "Authorization: Bearer $LAYERCLOUD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "laya",
    "messages": [{"role": "user", "content": "{\"state\": \"I was charged twice\", \"questions\": {\"department\": {\"type\": \"choice\", \"instructions\": \"Which team?\", \"criteria\": [\"billing\", \"technical\"]}}}"}]
  }'
```

Both return the same decision. The OpenAI shape **carries** a decision request in its message content rather than a prompt; free-form text is refused with a clear error, because a classifier asked a question nobody typed still returns a confident answer.

## Access

`laya` behaves like any other model. It appears in your model list, and your account's existing model permissions decide whether you can use it. There is no separate sign-up.

## Two things to know before you rely on it

**Confidence values are not calibrated.** A `0.98` does not mean "98% likely to be right". Treat the numbers as relative scores between the options you gave it, and read the `caveat` the response carries.

**It runs locally, on our own hardware.** That is why it is cheap and why your text does not leave the building for this particular call.


## The three question types

Every question you ask has a **type**, and the type is what makes the answer usable in code rather than something you have to parse again.

**`choice`** picks one label from a list you supply. Use it for routing and categorisation.

```json
{"department": {"type": "choice",
                "instructions": "Which team should handle this?",
                "criteria": ["billing", "technical", "sales", "other"]}}
```

**`score`** rates the text against levels you describe, **in order**, from lowest to highest. Use it for any judgement where you need a severity or a quality, not a category.

```json
{"urgency": {"type": "score",
             "instructions": "How urgent is this?",
             "criteria": ["can wait", "should be handled today", "service is down"]}}
```

**`noul`** lets the model abstain. This matters more than it looks: on real traffic some inputs genuinely do not fit any option, and a model forced to choose will pick something anyway. Giving it a way out is how you find out that your own categories are wrong.

```json
{"is_feedback": {"type": "noul", "instructions": "Is this product feedback rather than a request?"}}
```

**What the reply looks like**

```json
{
  "answers": {
    "department": {
      "type": "choice",
      "choice": "billing",
      "probabilities": {"billing": 0.9867, "technical": 0.0081, "sales": 0.0026, "other": 0.0026},
      "confidence": 0.9277
    }
  },
  "routing": {"model": "english", "reason": "English Latin text"},
  "confidence_calibrated": false,
  "caveat": "..."
}
```

You get the chosen label, the scores behind it, and a flag telling you the scores are not calibrated. Read that flag.

## How companies can use it

Laya is most useful where a **person is currently reading something short and deciding where it goes**. Those decisions are cheap individually and expensive in volume.

**Support triage.** Route an inbound ticket to the right queue before a human opens it, and let the `noul` type flag the ones that fit nothing - which are usually the interesting ones. The confidence scores let you automate the easy 80% and escalate the rest instead of pretending the whole stream is one difficulty.

**Lead qualification.** Score an enquiry form against levels you define, so sales sees a ranked list rather than a mailbox. The score is relative to the levels you wrote, so write them carefully - the quality of the answer is bounded by the quality of the options.

**Content moderation.** Classify a report against your own policy categories instead of a generic "toxic / not toxic". Because you supply the labels, the output matches the rules your team actually enforces, and the reasoning is auditable.

**Form and document understanding.** Where a field arrives as free text but must become a value, this turns prose into a typed answer: which department, how urgent, which product line, whether it is a complaint or a question.

**Language routing.** The service reports which model it used and why, so you can see whether a request was handled by the English or the multilingual path.

## What it is not for

It will not draft the reply. It will not summarise a long document into a report. It will not hold a conversation, and it does not remember anything between calls - each request is independent, so anything the model needs to know must be in the `state` you send.

If your problem is "read this and decide one thing", this is the right shape. If it is "write something for me", it is the wrong tool and a chat model is the right one.

## Cost and privacy

It runs locally on our own hardware. Your text does not leave for this call, and there is no per-token charge the way a hosted model has. The trade is that it is a small specialised model: excellent at the narrow job above, and not a general assistant.


## How fast is it

Two numbers matter, and they belong to different hardware.

**On our service, a short decision takes about one second.** That is our own measurement on our own CPU, over several calls, and it includes the network round trip through the gateway.

**The model's published figure is sub-35 ms**, measured on a T4 GPU by the people who built it. That is the same model on different hardware - a GPU is simply the right place to run a forward pass.

**What the gateway itself costs is about 25 ms.** Calling the service directly takes 1038 ms; calling it through the gateway takes 1062 ms. So routing, authentication, usage logging and the OpenAI-compatible layer together add roughly 25 milliseconds. **The time is the model and the hardware, not the plumbing.**

That is worth stating precisely because it tells you what to expect: if you are deciding where a support ticket should go, one second is nothing next to a person opening it. If you need the 35 ms figure, that is a GPU conversation.

## Watch

These are independent explainers of the Laya family of models - **made by other people, not by us.** We are linking them because they explain the technology better and faster than another page of prose would, not because they are our material.

- **[Laya: Open-Source Counterpart to Jev | Multilingual](https://www.youtube.com/watch?v=UoRNo5LoDRE)** - Mohamed Naji Aboo
- **[Laya: Non-Autoregressive Decision Model with RL](https://www.youtube.com/watch?v=5vIK_4yFpvY)** - Cold Boot
- **[Jev and Laya Explained: Decision Models vs LLMs](https://www.youtube.com/watch?v=fGl7T88bITo)** - TerraNet Technologies LLC
- **[Jev vs Laya: Closed or Open AI Decision Model?](https://www.youtube.com/watch?v=Kp22EHJpnts)** - Join dev

If you want the source of truth rather than a video, the project's own site is [laya.convaiinnovations.com](https://laya.convaiinnovations.com/) and the weights are published openly on Hugging Face.

**One thing to keep in mind while you watch.** Those videos describe the model. What we run is the same model behind our own gateway, on CPU, with an OpenAI-compatible front door and your existing API key. The *what* is identical; the *where* is not.

