> ## Documentation Index
> Fetch the complete documentation index at: https://docs.errorbar.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# List failure clusters

> See live production failures grouped into systemic causes per criterion (judge FAIL rationales plus pending scan suspects, clustered by a model) so a customer can find what to fix first rather than reading failures one by one.

400 "window_days must be an integer 1..90" for an out-of-range value. MONEY: a fresh clustering (cache miss or force=true) makes one small metered model call per criterion that has ≥4 failure reasons (at most 40 reasons per criterion) — billed to the wallet like other assists; cached responses cost nothing. Criteria with fewer than 4 reasons are listed with no clusters.



## OpenAPI

````yaml /openapi.json get /v1/evals/failure_clusters
openapi: 3.1.0
info:
  title: errorbar Management API
  description: >-
    The management API behind the improvement loop: capture and setup, request
    logs and datasets, grades (labels), judges (criteria), evals and deploy
    gates, fine-tuning and reinforcement learning, dedicated GPU endpoints, and
    model aliases and versions. Authenticated with a workspace API key
    (sk_sovereign_...). The inference API (chat, embeddings, rerank, responses)
    is OpenAI-compatible and documented separately.


    Responses are snake_case, list endpoints on the loop products use the
    {"object": "list", "data": [...]} envelope, and refusals use the same nested
    error shape the gateway emits: {"error": {"message", "type", "code"}}.
    Request bodies on the loop products (logs, labels, criteria, evals,
    datasets, aliases) are snake_case; the training and infrastructure products
    (fine-tuning, GRPO, environment tools, dedicated, model-version adoption)
    validate camelCase bodies, and each schema below says which it is. Endpoints
    that spend money require a key minted by a workspace owner or admin and
    return 403 otherwise.
  version: 1.0.0
servers:
  - url: https://gateway.errorbar.ai
    description: Production
  - url: https://www.errorbar.ai/api
    description: Control plane (also served at this base URL)
security:
  - bearerAuth: []
paths:
  /v1/evals/failure_clusters:
    get:
      tags:
        - Evals
      summary: List failure clusters
      description: >-
        See live production failures grouped into systemic causes per criterion
        (judge FAIL rationales plus pending scan suspects, clustered by a model)
        so a customer can find what to fix first rather than reading failures
        one by one.


        400 "window_days must be an integer 1..90" for an out-of-range value.
        MONEY: a fresh clustering (cache miss or force=true) makes one small
        metered model call per criterion that has ≥4 failure reasons (at most 40
        reasons per criterion) — billed to the wallet like other assists; cached
        responses cost nothing. Criteria with fewer than 4 reasons are listed
        with no clusters.
      operationId: getFailureClusters
      parameters:
        - name: window_days
          in: query
          schema:
            type: integer
          description: Look-back window in days, integer 1..90.
        - name: force
          in: query
          schema:
            type: boolean
          description: >-
            Pass the literal "true" to bypass the per-workspace one-hour cache
            and re-cluster now.
      responses:
        '200':
          description: >-
            {window_days, generated_at, cached (true when served from the hourly
            cache), criteria:[{criterion_id, criterion_name, failures (online
            FAILs + pending suspects, deduped), without_reason (failures with no
            stored rationale — counted, never clustered), clusters:[{name,
            count, share (of this criterion's clustered failures), request_ids,
            example (one representative rationale verbatim)}]}]}. Cache-Control:
            no-store.
          content:
            application/json:
              schema:
                type: object
        '400':
          $ref: '#/components/responses/BadRequest'
        '401':
          $ref: '#/components/responses/Unauthorized'
components:
  responses:
    BadRequest:
      description: Malformed request or invalid field.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
    Unauthorized:
      description: Missing, malformed, or revoked API key.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            error:
              message: Invalid API key
              type: invalid_request_error
              code: invalid_api_key
  schemas:
    Error:
      type: object
      properties:
        error:
          type: object
          properties:
            message:
              type: string
            type:
              type: string
              description: >-
                invalid_request_error, insufficient_quota, rate_limit_error, or
                api_error.
            code:
              type: string
              description: >-
                Machine-stable cause, e.g. invalid_api_key, not_found,
                insufficient_permissions, precondition_failed.
          required:
            - message
            - type
            - code
      description: >-
        Every refusal — gateway and management API alike — uses this one
        envelope.
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: >-
        Your workspace API key, e.g. `sk_sovereign_...`, sent as `Authorization:
        Bearer <key>`.

````