SQL Query Equivalence

import langwatch

df = langwatch.datasets.get_dataset("dataset-id").to_pandas()

experiment = langwatch.experiment.init("my-experiment")

for index, row in experiment.loop(df.iterrows()):
    # your execution code here
    experiment.evaluate(
        "ragas/sql_query_equivalence",
        index=index,
        data={
            "output": output,
            "expected_output": row["expected_output"],
            "expected_contexts": row["expected_contexts"],
        },
        settings={}
    )

[
  {
    "status": "processed",
    "score": 123,
    "passed": true,
    "label": "<string>",
    "details": "<string>",
    "cost": {
      "currency": "<string>",
      "amount": 123
    }
  }
]

POST

ragas

sql_query_equivalence

evaluate

import langwatch

df = langwatch.datasets.get_dataset("dataset-id").to_pandas()

experiment = langwatch.experiment.init("my-experiment")

for index, row in experiment.loop(df.iterrows()):
    # your execution code here
    experiment.evaluate(
        "ragas/sql_query_equivalence",
        index=index,
        data={
            "output": output,
            "expected_output": row["expected_output"],
            "expected_contexts": row["expected_contexts"],
        },
        settings={}
    )

[
  {
    "status": "processed",
    "score": 123,
    "passed": true,
    "label": "<string>",
    "details": "<string>",
    "cost": {
      "currency": "<string>",
      "amount": 123
    }
  }
]

Authorizations

X-Auth-Token

string

header

required

API key for authentication

Body

application/json

output

string

required

The output/response text to evaluate

expected_output

string

required

The expected output for comparison

expected_contexts

string[]

required

The expected contexts for comparison

settings

object

Show child attributes

Response

Successful evaluation

status

enum<string>

Available options:

processed,

skipped,

error

score

number

Numeric score from the evaluation

passed

boolean

Whether the evaluation passed

label

string

Label assigned by the evaluation

details

string

Additional details about the evaluation

cost

object

Show child attributes

ROUGE Score LLM-as-a-Judge Boolean Evaluator

⌘I

Traces

Prompts

Annotations

Datasets

Triggers

Scenarios

Evaluators

SQL Query Equivalence

Authorizations

Body

Response