convaiinnovationsOpen weights model

Laya-typed benchmarks

All models
  • General rank
    #19of 27Leads 0 of 26 (0%)non-coding tasks
  • Speed#3 / 8
    53 msMedian latency, p50Laptop GPU only
  • Cost
    Self-hostedRan on Laptop GPUNo hosted price
  • Search relevance#2 / 8 tied
    29.8%Best task · Accuracyvs small open models

Comparison summary

Strongest
General decisions
Leads 0 of 4 (0%)
Weakest
Agents
Leads 0 of 2 (0%)

Strongest at general decisions, weaker at agent actions. It leads 0 of 26 (0%) non-coding tasks, #19 of 27 models by that count.

Specifications

Lab
convaiinnovations
Weights
Open weights
Parameters
Not disclosed
Licence
Not recorded

Overall standing

Share of non-coding tasks led · #19 of 27

Safety and guardrails

Share of tasks led · #10 of 13

  1. Eikos-27BLeads 3 of 3
  2. AutoJev-27BLeads 2 of 2
  3. JevLeads 8 of 12
  4. Shisa DE-1Leads 1 of 4
  5. Kev-9BLeads 0 of 4
  6. Eikos-4BLeads 0 of 3
  7. SemIf (Qwen3.5-4B)Leads 0 of 4
  8. Laya-typed
    Leads 0 of 8
See the 4 tasks

Prompt injection

Tier 4 of 5#11 / 12

Catch rate at 5% false alarms, % · Higher is better

  1. Jev94.6, interval 92.2 to 96.7, tied for the lead
  2. Eikos-27B94.2, interval 90.5 to 97.1, tied for the lead
  3. Shisa DE-194.0, interval 90.3 to 96.2, tied for the lead
  4. AutoJev-27B92.0, interval 86.6 to 95.0, tied for the lead
  5. Kev-4B91.6, interval 88.2 to 94.2, tier 2
  6. Eikos-4B90.8, interval 88.3 to 94.0, tier 2
  7. SemIf (Qwen3.5-4B)90.6, interval 87.8 to 93.5, tier 2
  8. Laya-typed
    49.2, interval 42.6 to 58.2, tier 4
Tested against Jev and 14 open modelsAlso tested against small open models: #2 of 4 tied

Moderation

Tier 5 of 6#9 / 12

AUROC, % · Higher is better

  1. AutoJev-27B*90.6*, interval 88.4 to 92.7, trained on this data, not ranked
  2. Eikos-27B90.3, interval 87.9 to 92.4, tied for the lead
  3. Jev90.3, interval 87.9 to 92.4, tied for the lead
  4. Eikos-4B88.8, interval 86.3 to 91.3, tier 2
  5. SemIf (Qwen3.5-4B)88.7, interval 86.2 to 91.2, tier 2
  6. Kev-9B88.5, interval 86.0 to 90.9, tier 2
  7. Shisa DE-188.3, interval 85.6 to 90.8, tier 2
  8. Laya-typed
    70.0, interval 66.0 to 73.7, tier 5
Tested against Jev and 14 open modelsAlso tested against small open models: #3 of 5 tied

Off-topic

Tier 8 of 8#8 / 9

Balanced accuracy, % · Higher is better

  1. Jev93.4, interval 91.8 to 94.9, ahead alone
  2. AutoJev-27B*92.7*, interval 91.0 to 94.4, trained on this data, not ranked
  3. Eikos-27B*90.6*, interval 88.8 to 92.4, trained on this data, not ranked
  4. Kev-9B89.6, interval 87.7 to 91.5, tier 2
  5. Shisa DE-188.0, interval 86.1 to 89.9, tier 3
  6. Kev-4B84.4, interval 82.1 to 86.7, tier 4
  7. SemIf (Qwen3.5-4B)81.1, interval 79.0 to 83.4, tier 5
  8. Laya-typed
    52.6, interval 49.6 to 55.6, tier 8
Tested against Jev and 14 open modelsAlso tested against small open models: #4 of 5 tied

PII

Tier 5 of 5#8 / 13

Catch rate at 5% false alarms, % · Higher is better

  1. Eikos-27B95.2, interval 92.8 to 97.8, tied for the lead
  2. AutoJev-27B92.8, interval 90.2 to 95.5, tied for the lead
  3. Jev90.8, interval 88.1 to 93.8, tier 2
  4. Shisa DE-184.2, interval 77.3 to 89.1, tier 3
  5. Eikos-4B78.6, interval 74.5 to 85.3, tier 3
  6. SemIf (Qwen3.5-4B)66.1, interval 60.9 to 74.6, tier 4
  7. Kev-9B62.5, interval 0.2 to 66.4, tier 4
  8. Laya-typed
    21.0, interval 14.6 to 29.1, tier 5
Tested against Jev and 14 open modelsAlso tested against small open models: #2 of 5 tied

Routing and classification

Share of tasks led · #10 of 16

  1. AutoJev-27BLeads 4 of 5
  2. Shisa DE-1Leads 3 of 4
  3. JevLeads 9 of 14
  4. Eikos-27BLeads 1 of 3
  5. SemIf (Qwen3.5-4B)Leads 1 of 3
  6. Kev-4BLeads 0 of 3
  7. Kev-9BLeads 0 of 2
  8. Laya-typed
    Leads 0 of 8
See the 5 tasks

Routing, 20 intents

Tier 4 of 5#5 / 7

Accuracy, % · Higher is better

  1. Kev-4B*93.0*, interval 91.4 to 94.5, trained on this data, not ranked
  2. Kev-9B*93.0*, interval 91.4 to 94.6, trained on this data, not ranked
  3. Kev-0.8B*91.3*, interval 89.5 to 93.0, trained on this data, not ranked
  4. AutoJev-27B89.7, interval 87.7 to 91.6, ahead alone
  5. Eikos-27B*89.6*, interval 87.7 to 91.4, trained on this data, not ranked
  6. Kev-0.6B*89.5*, interval 87.5 to 91.5, trained on this data, not ranked
  7. Jev89.1, interval 87.0 to 91.0, tier 2
  8. Laya-typed
    77.0, interval 74.2 to 79.5, tier 4
Tested against Jev and 14 open modelsAlso tested against small open models: #2 of 4 tied

Routing, 77 intents

Tier 3 of 3#4 / 5

Accuracy, % · Higher is better

  1. Kev-4B*85.0*, interval 81.8 to 88.0, trained on this data, not ranked
  2. Kev-9B*85.0*, interval 81.8 to 88.2, trained on this data, not ranked
  3. Kev-0.8B*83.0*, interval 79.6 to 86.4, trained on this data, not ranked
  4. AutoJev-27B80.6, interval 77.0 to 84.0, tied for the lead
  5. Jev79.6, interval 76.0 to 83.0, tied for the lead
  6. Kev-0.6B*78.2*, interval 74.4 to 81.8, trained on this data, not ranked
  7. Eikos-27B*77.8*, interval 74.0 to 81.4, trained on this data, not ranked
  8. Laya-typed
    40.0, interval 35.8 to 44.4, tier 3
Tested against Jev and 14 open modelsAlso tested against small open models: #2 of 3 tied

Tool routing

Tier 9 of 10#14 / 16

Accuracy, % · Higher is better

  1. Shisa DE-181.3, interval 78.9 to 83.5, tied for the lead
  2. AutoJev-27B80.6, interval 78.1 to 82.9, tied for the lead
  3. SemIf (Qwen3.5-4B)79.9, interval 77.4 to 82.2, tied for the lead
  4. Eikos-27B79.3, interval 76.8 to 81.8, tier 2
  5. Jev78.3, interval 75.7 to 80.7, tier 2
  6. Eikos-4B75.5, interval 72.8 to 78.0, tier 3
  7. Kev-9B72.4, interval 69.7 to 75.0, tier 4
  8. Laya-typed
    21.5, interval 19.0 to 24.0, tier 9
Tested against Jev and 14 open modelsAlso tested against small open models: #6 of 8 tied

Complaint routing

Tier 6 of 8#10 / 15

Accuracy, % · Higher is better

  1. Jev78.7, interval 76.3 to 81.2, tied for the lead
  2. Shisa DE-178.0, interval 75.4 to 80.6, tied for the lead
  3. Kev-4B76.2, interval 73.6 to 79.0, tier 2
  4. AutoJev-27B75.9, interval 73.2 to 78.7, tier 2
  5. Eikos-27B75.6, interval 73.0 to 78.4, tier 2
  6. Kev-9B75.0, interval 72.4 to 77.7, tier 2
  7. SemIf (Qwen3.5-4B)72.9, interval 70.2 to 75.7, tier 3
  8. Laya-typed
    47.1, interval 44.1 to 50.1, tier 6
Tested against Jev and 13 open modelsAlso tested against small open models: #3 of 8 tied

Typed decisions

Trained on the test data, not ranked

Accuracy, % · Higher is better

  1. Laya-typed*
    77.4*, interval 75.4 to 79.4, trained on this data, not ranked
  2. Shisa DE-175.1, interval 72.8 to 77.3, tied for the lead
  3. Jev73.9, interval 71.8 to 76.0, tied for the lead
  4. AutoJev-27B73.8, interval 71.7 to 75.9, tied for the lead
  5. Eikos-27B73.4, interval 71.2 to 75.6, tied for the lead
  6. Kev-4B65.8, interval 63.6 to 68.0, tier 2
  7. Eikos-4B65.3, interval 62.7 to 68.0, tier 2
  8. SemIf (Qwen3.5-4B)63.0, interval 60.5 to 65.4, tier 3
Tested against Jev and 13 open modelsAlso tested against small open models: trained on the data, not ranked

Grounding

Share of tasks led · #11 of 15

  1. Eikos-27BLeads 2 of 2
  2. JevLeads 4 of 6
  3. AutoJev-27BLeads 0 of 2
  4. Shisa DE-1Leads 0 of 2
  5. Eikos-4BLeads 0 of 2
  6. Kev-4BLeads 0 of 2
  7. Kev-9BLeads 0 of 2
  8. Laya-typed
    Leads 0 of 4
See the 2 tasks

RAG faithfulness

Tier 5 of 5#11 / 13

Balanced accuracy, % · Higher is better

  1. Eikos-27B82.1, interval 79.2 to 84.9, tied for the lead
  2. Jev80.3, interval 77.5 to 83.1, tied for the lead
  3. Shisa DE-179.2, interval 76.2 to 82.2, tier 2
  4. AutoJev-27B79.1, interval 76.1 to 82.0, tier 2
  5. Kev-9B73.1, interval 70.0 to 76.0, tier 3
  6. Kev-4B72.5, interval 69.4 to 75.5, tier 3
  7. SemIf (Qwen3.5-4B)72.0, interval 68.9 to 75.2, tier 3
  8. Laya-typed
    49.4, interval 48.2 to 50.6, tier 5
Tested against Jev and 14 open modelsAlso tested against small open models: #5 of 6 tied

Search relevance

Tier 5 of 7#9 / 16

Accuracy, % · Higher is better

  1. Jev57.7, interval 54.7 to 60.8, tied for the lead
  2. Eikos-27B56.6, interval 53.4 to 59.7, tied for the lead
  3. AutoJev-27B52.8, interval 49.6 to 55.9, tier 2
  4. Shisa DE-150.3, interval 47.1 to 53.6, tier 2
  5. Kev-4B44.8, interval 41.7 to 48.0, tier 3
  6. Kev-9B43.4, interval 40.2 to 46.5, tier 3
  7. Eikos-4B43.3, interval 40.1 to 46.3, tier 3
  8. Laya-typed
    29.8, interval 27.0 to 32.6, tier 5
Tested against Jev and 14 open modelsAlso tested against small open models: #2 of 8 tied

Agents

Share of tasks led · #6 of 8

  1. JevLeads 2 of 2
  2. SimpleJev (Qwen3.5-0.8B)Leads 0 of 2
  3. Kev-0.8BLeads 0 of 2
  4. Kev-0.6BLeads 0 of 2
  5. LayaLeads 0 of 2
  6. Laya-typed
    Leads 0 of 2
  7. openJev Verdict 1.4Leads 0 of 2
  8. SemIf (Qwen3-0.6B)Leads 0 of 2
See the task

Web-agent actions

Tier 10 of 10#12 / 15

Accuracy, % · Higher is better

  1. Jev70.8, interval 68.9 to 72.6, ahead alone
  2. Eikos-27B68.4, interval 66.5 to 70.4, tier 2
  3. AutoJev-27B68.2, interval 66.2 to 70.2, tier 2
  4. Shisa DE-165.2, interval 63.3 to 67.2, tier 3
  5. Eikos-4B62.9, interval 60.9 to 64.8, tier 4
  6. SemIf (Qwen3.5-4B)59.5, interval 57.5 to 61.5, tier 5
  7. Kev-4B58.0, interval 55.9 to 60.2, tier 5
  8. Laya-typed
    18.3, interval 16.7 to 19.9, tier 10
Tested against Jev and 13 open modelsAlso tested against small open models: #5 of 8 tied

General decisions

Share of tasks led · #6 of 10

  1. JevLeads 4 of 4
  2. AutoJev-27BLeads 2 of 2
  3. Kev-0.6BLeads 0 of 4
  4. Kev-0.8BLeads 0 of 4
  5. LayaLeads 0 of 4
  6. Laya-typed
    Leads 0 of 4
  7. SimpleJev (Qwen3.5-0.8B)Leads 0 of 2
  8. Kev-4BLeads 0 of 2
See the 2 tasks

Community sets

Tier 4 of 5#6 / 10

Accuracy, % · Higher is better

  1. AutoJev-27B63.0, interval 60.2 to 65.8, tied for the lead
  2. Jev62.3, interval 59.5 to 65.1, tied for the lead
  3. Eikos-27B59.8, interval 56.9 to 62.6, tier 2
  4. Eikos-4B55.8, interval 52.9 to 58.7, tier 3
  5. Kev-4B55.8, interval 53.0 to 58.7, tier 3
  6. Laya-typed
    47.8, interval 44.9 to 50.5, tier 4
  7. Kev-0.8B46.4, interval 43.5 to 49.3, tier 4
  8. Laya45.9, interval 43.1 to 48.8, tier 4
Tested against Jev and 13 open modelsAlso tested against small open models: #2 of 5 tied

JevBench public

Tier 4 of 5#6 / 12

Accuracy, % · Higher is better

  1. Eikos-27B*91.8*, interval 87.9 to 95.2, trained on this data, not ranked
  2. AutoJev-27B86.6, interval 81.9 to 90.9, tied for the lead
  3. Shisa DE-186.1, interval 81.2 to 90.6, tied for the lead
  4. Jev85.7, interval 81.0 to 90.1, tied for the lead
  5. Eikos-4B*85.3*, interval 80.4 to 89.8, trained on this data, not ranked
  6. SemIf (Qwen3.5-4B)80.1, interval 74.7 to 85.3, tier 2
  7. Kev-4B71.4, interval 65.2 to 77.4, tier 3
  8. Laya-typed
    53.7, interval 46.9 to 60.6, tier 4
Tested against Jev and 13 open modelsAlso tested against small open models: #2 of 7 tied

Speed

Median latency, p50 · Laptop GPU only · lower is better

How we count

A model leads a task when it is in the task's leading tie tier. Each task counts inside its own benchmark, against that benchmark's models; no score is averaged. An area chart shows the models ranked on at least 2 of the area's tasks in a benchmark this model is in too. The overall standing counts every non-coding tasks and needs 5 for a rank. Latency is compared only on one hardware tier.

Laya-typedOther modelsTrained on the test data*Whisker: 95% interval