Playground: sinta o que um modelo de juízo faz

Com a mesma entrada: à esquerda, a saída típica de um modelo generativo; à direita, o juízo tipado do Jev. Todas as respostas são resultados reais da API pré-gravados (jev-1.13.0) com o uso real de tokens. Edite a entrada — a resposta permanece fixa, pertence à entrada original gravada.

Primitiva · Noul

Esta mensagem de suporte é urgente?

Noul avalia uma pergunta de sim/não e devolve a probabilidade de «sim» — um valor de 0 a 1.

questions (o pedido)
{
  "urgency": {
    "type": "noul",
    "instructions": "Does this message express urgency?"
  }
}
Um modelo generativo produziria

I'm really sorry to hear you've been having trouble connecting your Stripe account. I understand how frustrating that must be, especially when it affects your sales. Let me help you look into this. Could you tell me a bit more about...

Jev devolvePré-gravado · não é inferência ao vivo
noul = 0.98
  • Sim · urgente0.98
{
  "model": "jev-1.13.0",
  "answers": {
    "urgency": {
      "type": "noul",
      "noul": 0.98
    }
  }
}
Esta chamada: 301 tokens de entrada (a saída é gratuita) · 21 output
Primitiva · Choice

O que este e-mail quer?

Choice escolhe uma opção de um conjunto fixo, com probabilidades por opção e uma confiança geral.

questions (o pedido)
{
  "intent": {
    "type": "choice",
    "instructions": "What is the primary intent of this message?",
    "criteria": {
      "bug_report": "Reporting something broken",
      "feature_request": "Asking for new functionality",
      "how_to_question": "Asking how to do something",
      "other": "None of the above"
    }
  }
}
Um modelo generativo produziria

Thanks for reaching out! It sounds like you are experiencing an error with the export functionality on the reports page when using Safari. This could be related to browser compatibility. Have you tried...

Jev devolvePré-gravado · não é inferência ao vivo
choice = bug_report · confiança 1.0
  • Relatório de erro1.00
  • Pergunta de uso0.02
  • Outro0.01
{
  "model": "jev-1.13.0",
  "answers": {
    "intent": {
      "type": "choice",
      "choice": "bug_report",
      "confidence": 1,
      "probabilities": {
        "bug_report": 1,
        "feature_request": 0,
        "how_to_question": 0,
        "other": 0
      }
    }
  }
}
Esta chamada: 389 tokens de entrada (a saída é gratuita) · 50 output
Primitiva · Score

Quão grave é este problema?

Score avalia segundo critérios ordenados; o valor pode ficar entre níveis, com probabilidades por nível e confiança.

questions (o pedido)
{
  "severity": {
    "type": "score",
    "instructions": "How severe is the issue?",
    "criteria": [
      "Minor",
      "Moderate",
      "Critical"
    ]
  }
}
Um modelo generativo produziria

I'm so sorry to hear that your payouts have been failing for three days. That sounds really stressful. Payout failures can have several causes, such as... (510 more words)

Jev devolvePré-gravado · não é inferência ao vivo
score = 1,87 (entre Moderado e Crítico) · confiança 0,81
  • Leve0.00
  • Moderado0.13
  • Crítico0.87
{
  "model": "jev-1.13.0",
  "answers": {
    "severity": {
      "type": "score",
      "score": 1.87,
      "confidence": 0.81,
      "legend": {
        "0": "Minor",
        "1": "Moderate",
        "2": "Critical"
      },
      "probabilities": {
        "0": 0,
        "1": 0.13,
        "2": 0.87
      }
    }
  }
}
Esta chamada: 321 tokens de entrada (a saída é gratuita) · 36 output
Padrão de pipeline

Pontuação composta: responder com que rapidez?

Combina duas pontuações independentes —intensidade emocional e impacto de negócio— numa prioridade de resposta ponderada. Tipado e explicável.

questions (o pedido)
{
  "emotional_intensity": {
    "type": "score",
    "instructions": "How emotionally intense is this message?",
    "criteria": [
      "Calm and factual",
      "Frustrated but professional",
      "Angry or escalating"
    ]
  },
  "business_impact": {
    "type": "score",
    "instructions": "How severe is the business impact described?",
    "criteria": [
      "No material impact",
      "Degraded operations",
      "Revenue-blocking outage"
    ]
  }
}
Um modelo generativo produziria

Dear valued customer, thank you for bringing this to our attention. We sincerely apologize for the inconvenience caused by the API errors. Our engineering team is investigating the issue with the highest priority...

Jev devolvePré-gravado · não é inferência ao vivo
priority = 0.6×2.0 + 0.4×1.82 = 1.93 → responder imediatamente
  • Emoção · escalada0.82
  • Impacto · bloqueio de receita1.00
{
  "model": "jev-1.13.0",
  "answers": {
    "emotional_intensity": {
      "type": "score",
      "score": 1.82,
      "confidence": 0.74,
      "legend": {
        "0": "Calm and factual",
        "1": "Frustrated but professional",
        "2": "Angry or escalating"
      },
      "probabilities": {
        "0": 0,
        "1": 0.18,
        "2": 0.82
      }
    },
    "business_impact": {
      "type": "score",
      "score": 2,
      "confidence": 1,
      "legend": {
        "0": "No material impact",
        "1": "Degraded operations",
        "2": "Revenue-blocking outage"
      },
      "probabilities": {
        "0": 0,
        "1": 0,
        "2": 1
      }
    }
  }
}
Esta chamada: 397 tokens de entrada (a saída é gratuita) · 37 output
Padrão de pipeline

Fan-out: todas as decisões numa chamada

Sem «classificar e depois perguntar». Envie de uma vez todas as perguntas que possa precisar — as respostas são independentes e o seu código decide quais usar. Tokens de saída são gratuitos.

questions (o pedido)
{
  "request_type": {
    "type": "choice",
    "instructions": "What kind of request is this?",
    "criteria": {
      "sales_question": "Asking about plans, pricing, compliance, or contracts",
      "billing_issue": "Payment, invoice, or refund matters",
      "technical_support": "Something is broken or needs troubleshooting"
    }
  },
  "is_existing_customer": {
    "type": "noul",
    "instructions": "Does the writer appear to be an existing customer?"
  },
  "tone_politeness": {
    "type": "score",
    "instructions": "How courteous is the tone overall?",
    "criteria": [
      "Blunt or rude",
      "Neutral",
      "Warm and appreciative"
    ]
  }
}
Um modelo generativo produziria

Thank you for reaching out! To answer your questions: regarding SOC 2 report access, our enterprise plan does include... (the answer continues for 300+ words, mixing both topics with no typed structure to route on)

Jev devolvePré-gravado · não é inferência ao vivo
choice=sales_question (conf 0.25) · noul=0.86 · score=1.74 — uma chamada, três decisões independentes
  • Tipo · vendas 0.500.50
  • Cliente existente · sim 0.860.86
  • Tom · caloroso 0.740.74
{
  "model": "jev-1.13.0",
  "answers": {
    "request_type": {
      "type": "choice",
      "choice": "sales_question",
      "confidence": 0.25,
      "probabilities": {
        "sales_question": 0.5,
        "billing_issue": 0.5,
        "technical_support": 0
      }
    },
    "is_existing_customer": {
      "type": "noul",
      "noul": 0.86
    },
    "tone_politeness": {
      "type": "score",
      "score": 1.74,
      "confidence": 0.61,
      "legend": {
        "0": "Blunt or rude",
        "1": "Neutral",
        "2": "Warm and appreciative"
      },
      "probabilities": {
        "0": 0,
        "1": 0.26,
        "2": 0.74
      }
    }
  }
}
Esta chamada: 459 tokens de entrada (a saída é gratuita) · 78 output