Reports Module

The reports module generates evaluation reports in various formats.

JSON Reports

toolscore.reports.generate_json_report(result, output_path='toolscore.json')[source]

Generate JSON report from evaluation result.

Parameters:
  • result (EvaluationResult) – Evaluation result to report.

  • output_path (str | Path) – Path to save the JSON report.

Return type:

Path

Returns:

Path to the generated report file.

HTML Reports

toolscore.reports.generate_html_report(result, output_path='toolscore.html')[source]

Generate HTML report from evaluation result.

Parameters:
  • result (EvaluationResult) – Evaluation result to report.

  • output_path (str | Path) – Path to save the HTML report.

Return type:

Path

Returns:

Path to the generated report file.

Markdown and CSV Reports

toolscore.reports.generate_markdown_report(result, output_path='toolscore.md')[source]

Generate Markdown report from evaluation result.

Creates a Markdown file perfect for embedding in GitHub issues, PRs, wikis, or documentation. Includes formatted tables and emoji indicators.

Parameters:
  • result (EvaluationResult) – Evaluation result to report.

  • output_path (str | Path) – Path to save the Markdown report.

Return type:

Path

Returns:

Path to the generated report file.

toolscore.reports.generate_csv_report(result, output_path='toolscore.csv')[source]

Generate CSV report from evaluation result.

Creates a CSV file with metrics that can be opened in Excel or Google Sheets. Perfect for sharing results with non-technical stakeholders.

Parameters:
  • result (EvaluationResult) – Evaluation result to report.

  • output_path (str | Path) – Path to save the CSV report.

Return type:

Path

Returns:

Path to the generated report file.

Behavior and Safety Findings

toolscore.reports.findings.behavior_findings(result)[source]

Collect behavior and safety findings from an evaluation result.

Parameters:

result (EvaluationResult) – The evaluation result.

Return type:

list[tuple[str, str]]

Returns:

(severity, message) pairs, "error" first (forbidden calls and credentials), then "warning" (failed calls and blind retries). Call numbers are 1-based positions in the actual trace.