Playground: siente lo que hace un modelo de juicio

Con la misma entrada: a la izquierda, la salida típica de un modelo generativo; a la derecha, el juicio tipado de Jev. Todas las respuestas son resultados reales de la API pregrabados (jev-1.13.0) con el uso real de tokens. Edita la entrada — la respuesta no cambia: pertenece a la entrada original grabada.

Primitiva · Noul

¿Es urgente este mensaje de soporte?

Noul evalúa una pregunta de sí/no y devuelve la probabilidad de «sí»: un valor de 0 a 1.

questions (la solicitud)
{
  "urgency": {
    "type": "noul",
    "instructions": "Does this message express urgency?"
  }
}
Un modelo generativo produciría

I'm really sorry to hear you've been having trouble connecting your Stripe account. I understand how frustrating that must be, especially when it affects your sales. Let me help you look into this. Could you tell me a bit more about...

Jev devuelvePregrabado · no es inferencia en vivo
noul = 0.98
  • Sí · urgente0.98
{
  "model": "jev-1.13.0",
  "answers": {
    "urgency": {
      "type": "noul",
      "noul": 0.98
    }
  }
}
Esta llamada: 301 tokens de entrada (la salida es gratis) · 21 output
Primitiva · Choice

¿Qué quiere este correo?

Choice elige una opción de un conjunto fijo, con probabilidades por opción y una confianza general.

questions (la solicitud)
{
  "intent": {
    "type": "choice",
    "instructions": "What is the primary intent of this message?",
    "criteria": {
      "bug_report": "Reporting something broken",
      "feature_request": "Asking for new functionality",
      "how_to_question": "Asking how to do something",
      "other": "None of the above"
    }
  }
}
Un modelo generativo produciría

Thanks for reaching out! It sounds like you are experiencing an error with the export functionality on the reports page when using Safari. This could be related to browser compatibility. Have you tried...

Jev devuelvePregrabado · no es inferencia en vivo
choice = bug_report · confianza 0.97
  • Informe de error1.00
  • Pregunta de uso0.02
  • Otro0.01
{
  "model": "jev-1.13.0",
  "answers": {
    "intent": {
      "type": "choice",
      "choice": "bug_report",
      "confidence": 1,
      "probabilities": {
        "bug_report": 1,
        "feature_request": 0,
        "how_to_question": 0,
        "other": 0
      }
    }
  }
}
Esta llamada: 389 tokens de entrada (la salida es gratis) · 50 output
Primitiva · Score

¿Qué tan grave es este problema?

Score califica según criterios ordenados; el valor puede quedar entre niveles, con probabilidades por nivel y confianza.

questions (la solicitud)
{
  "severity": {
    "type": "score",
    "instructions": "How severe is the issue?",
    "criteria": [
      "Minor",
      "Moderate",
      "Critical"
    ]
  }
}
Un modelo generativo produciría

I'm so sorry to hear that your payouts have been failing for three days. That sounds really stressful. Payout failures can have several causes, such as... (510 more words)

Jev devuelvePregrabado · no es inferencia en vivo
score = 1,87 (entre Moderado y Crítico) · confianza 0,81
  • Leve0.00
  • Moderado0.13
  • Crítico0.87
{
  "model": "jev-1.13.0",
  "answers": {
    "severity": {
      "type": "score",
      "score": 1.87,
      "confidence": 0.81,
      "legend": {
        "0": "Minor",
        "1": "Moderate",
        "2": "Critical"
      },
      "probabilities": {
        "0": 0,
        "1": 0.13,
        "2": 0.87
      }
    }
  }
}
Esta llamada: 321 tokens de entrada (la salida es gratis) · 36 output
Patrón de pipeline

Puntuación compuesta: ¿con qué rapidez responder?

Combina dos puntuaciones independientes —intensidad emocional e impacto de negocio— en una prioridad de respuesta ponderada. Tipado y explicable.

questions (la solicitud)
{
  "emotional_intensity": {
    "type": "score",
    "instructions": "How emotionally intense is this message?",
    "criteria": [
      "Calm and factual",
      "Frustrated but professional",
      "Angry or escalating"
    ]
  },
  "business_impact": {
    "type": "score",
    "instructions": "How severe is the business impact described?",
    "criteria": [
      "No material impact",
      "Degraded operations",
      "Revenue-blocking outage"
    ]
  }
}
Un modelo generativo produciría

Dear valued customer, thank you for bringing this to our attention. We sincerely apologize for the inconvenience caused by the API errors. Our engineering team is investigating the issue with the highest priority...

Jev devuelvePregrabado · no es inferencia en vivo
priority = 0,6×2 + 0,4×1,82 = 1,93 → responder de inmediato
  • Emoción · escalada0.82
  • Impacto · bloqueo de ingresos1.00
{
  "model": "jev-1.13.0",
  "answers": {
    "emotional_intensity": {
      "type": "score",
      "score": 1.82,
      "confidence": 0.74,
      "legend": {
        "0": "Calm and factual",
        "1": "Frustrated but professional",
        "2": "Angry or escalating"
      },
      "probabilities": {
        "0": 0,
        "1": 0.18,
        "2": 0.82
      }
    },
    "business_impact": {
      "type": "score",
      "score": 2,
      "confidence": 1,
      "legend": {
        "0": "No material impact",
        "1": "Degraded operations",
        "2": "Revenue-blocking outage"
      },
      "probabilities": {
        "0": 0,
        "1": 0,
        "2": 1
      }
    }
  }
}
Esta llamada: 397 tokens de entrada (la salida es gratis) · 37 output
Patrón de pipeline

Fan-out: todas las decisiones en una llamada

Sin «clasificar y luego preguntar». Envía de una vez todas las preguntas que puedas necesitar — las respuestas son independientes y tu código decide cuáles usar. Los tokens de salida son gratis.

questions (la solicitud)
{
  "request_type": {
    "type": "choice",
    "instructions": "What kind of request is this?",
    "criteria": {
      "sales_question": "Asking about plans, pricing, compliance, or contracts",
      "billing_issue": "Payment, invoice, or refund matters",
      "technical_support": "Something is broken or needs troubleshooting"
    }
  },
  "is_existing_customer": {
    "type": "noul",
    "instructions": "Does the writer appear to be an existing customer?"
  },
  "tone_politeness": {
    "type": "score",
    "instructions": "How courteous is the tone overall?",
    "criteria": [
      "Blunt or rude",
      "Neutral",
      "Warm and appreciative"
    ]
  }
}
Un modelo generativo produciría

Thank you for reaching out! To answer your questions: regarding SOC 2 report access, our enterprise plan does include... (the answer continues for 300+ words, mixing both topics with no typed structure to route on)

Jev devuelvePregrabado · no es inferencia en vivo
choice=sales_question (conf 0.25) · noul=0,86 · score=1,74 — una llamada, tres decisiones independientes
  • Tipo · ventas 0,500.50
  • Cliente existente · sí 0,860.86
  • Tono · cálido 0,740.74
{
  "model": "jev-1.13.0",
  "answers": {
    "request_type": {
      "type": "choice",
      "choice": "sales_question",
      "confidence": 0.25,
      "probabilities": {
        "sales_question": 0.5,
        "billing_issue": 0.5,
        "technical_support": 0
      }
    },
    "is_existing_customer": {
      "type": "noul",
      "noul": 0.86
    },
    "tone_politeness": {
      "type": "score",
      "score": 1.74,
      "confidence": 0.61,
      "legend": {
        "0": "Blunt or rude",
        "1": "Neutral",
        "2": "Warm and appreciative"
      },
      "probabilities": {
        "0": 0,
        "1": 0.26,
        "2": 0.74
      }
    }
  }
}
Esta llamada: 459 tokens de entrada (la salida es gratis) · 78 output