Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, Laya, Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-9B, Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Eikos-4B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-9B, Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Eikos-4B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-9B, Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Eikos-4B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-9B, Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Eikos-4B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-9B, Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Eikos-4B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev, SemIf (Qwen3-0.6B), SemIf (Qwen3.5-4B).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, SemIf (Qwen3-0.6B), Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Eikos-4B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev, openJev Verdict 1.4, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), SemIf (Qwen3.5-4B), Shisa DE-1.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, openJev Verdict 1.4, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Shisa DE-1, Eikos-4B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-9B, Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Eikos-4B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-9B, Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-9B, Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Eikos-4B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-9B, Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Eikos-4B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, Kev-9B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, Kev-4B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Eikos-4B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, Kev-4B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Eikos-4B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, Kev-4B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Eikos-4B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate). Not plotted (no latency on NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate) measured): Jev, Kev-4B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA L4, 24 GB (cloud). Not plotted (no latency on NVIDIA L4, 24 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Shisa DE-1, AutoJev-27B, Eikos-27B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: NVIDIA H100 PCIe, 80 GB (cloud). Not plotted (no latency on NVIDIA H100 PCIe, 80 GB (cloud) measured): Jev, Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Eikos-4B, GLiNER2.5-Decide.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), Kev-4B, SemIf (Qwen3.5-4B), Shisa DE-1, AutoJev-27B, Eikos-4B, Eikos-27B, GLiNER2.5-Decide.