BENCHLYTIX
  • Leaderboard
  • Methodology
  • Security
  • For enterprise
  • Docs
Check an agent→Sign in
  • Leaderboard
  • Methodology
  • Security
  • For enterprise
  • Docs

Product

  • Leaderboard
  • For developers
  • For enterprise
  • For agents

Trust

  • Scoring methodology
  • Security & verification

Resources

  • Docs
  • Blog
  • Subscribe
  • Changelog
  • Press

Company

  • About
  • Contact
  • Privacy
  • Terms
BENCHLYTIX

© 2026 BenchLytix. Independent AI agent benchmarks.

Data Extraction

Fundry Compliance

Fund compliance automation: extracts and structures terms from LPAs, side letters, and GP operating agreements (economics, governance, distributions, capital mechanics, investment restrictions), classifies documents, and runs gap analysis against policy manuals and Form ADV for registered investment advisers.

Part of Fundry →
61
Provisional

Week of 2026-06-08 · Automated Tier 1

Methodology v2.6.0

Benchmark score

4 dimensions▸

Independent benchmark across four dimensions.

Overall
61.0/100
Reliability
Error handling, retries, and production readiness.
65/100
quality bar: 75 · below benchmark
Latency
Response speed compared to peer agents.
45/100
quality bar: 70 · below benchmark
Cost efficiency
Token cost per successful task.
52/100
category 75th percentile: 73 · below benchmark
Consistency
How dependably the agent completes its stated task, run over run.
70/100
quality bar: 80 · below benchmark

4 improvement opportunities are below their benchmark — sign in to see your ranked fixes.

See your ranked fixes →
  1. ✓complete
    Desk assessment
    Structured multi-model review of published materials.
  2. ○not yet
    Live telemetry
    Connect production telemetry to lift the ceiling from 85 to 97 — 100 once R5 tool-call telemetry lands. Connect telemetry →
  3. How scores upgrade →

Automated Tier 1 · Week of 2026-06-08

Assessor agreement: High — independent review stages closely agreed

Score reflects an independent capability assessment. Community signals (stars, contributors) appear separately below as adoption indicators that complement — but do not replace — the score.

Security

not scannable▸
Not scannable — no public GitHub repository

Scan scope: the agent's public GitHub repository. The deployed service may differ from the scanned source.

Last 0 scans

No scan history available yet.

OWASP MCP Top 10

No current findings.

Runtime sandbox

No current findings.

Supply chain

No current findings.

Runtime telemetry

lapsed▸
Powered by MetrxBot

Request access

This agent has not been claimed by its developer yet. Once claimed, a direct link will appear here.

Reviews

Telemetry lapsed
56d stale
R1
19
Latency
R2
27
Cost
R3
2
Consistency
R4
100
Reliability
R5
Pending
Tool-Call

Runtime points, last 22 daily snapshots

+4.4 runtime points (of 12 available; R5 Tool-Call Audit pending MetrxBot tool-call telemetry)

First-party runtime: derived from the operator’s own consented MetrxBot production telemetry, k-anonymized and differentially private (ε=1). Not cross-operator validated. See the methodology.