Evaluation Methodology
This page documents the full methodology behind the Global + China VC AI-Affinity Public Evidence Top 30: who evaluates, the evaluation agent pipeline, scoring dimensions, endpoint probe rules, the whitelist policy, the evidence policy, and known limitations.
Evidence collection, endpoint probes, and scoring for this index are performed entirely by autonomous AI agents. No per-firm human scoring is involved.
Every step of the evaluation—evidence collection, endpoint probes, per-dimension scoring, and ranking reconciliation—is performed automatically by AI agents. There is no per-firm human scoring or manual rank adjustment. When an institution requests a public-evidence re-review, the same agent pipeline reruns with the same rubric.
The evaluation agent pipeline
Four specialized agent roles produce the index in sequence, each with a bounded, auditable job:
- Source-discovery agents collect only first-party public evidence: official websites, official blogs and research, official open-source repositories, and machine-readable files published on the firm’s own site (llms.txt, Markdown pages, open metadata). Third-party media coverage is never used as scoring input.
- The endpoint-probe agent runs a bounded discovery pass against each primary domain with a 12-second timeout per request, recording status code, Content-Type, WWW-Authenticate header, final redirect URL, and MCP-related signals (such as WordPress namespace/route matches). Probing is fully scripted, and the raw per-firm response records remain auditable.
- Scoring agents grade each of the five dimensions independently against a rubric, from 0 up to that dimension’s weight, using only evidence available on the audit date.
- The reconciliation agent cross-checks evidence-to-score consistency, produces the ordering, applies the whitelist threshold, and pins the version and audit date; rerunning the same version must reproduce the same result.
Scoring dimensions and weights
The total score is the sum of five dimensions, out of 100.
| Dimension | Weight | Definition |
|---|---|---|
| MCP readiness | 25 | Verified public MCP/OAuth endpoints receive the highest credit; writing about MCP is scored separately. |
| Harness & agent fluency | 25 | Technical depth on agent harnesses, coding agents, evals, context engineering, and production architecture. |
| Operating adoption evidence | 20 | Public evidence that the firm or its programs use AI-native workflows, not merely invest in AI companies. |
| AI thesis & portfolio proof | 20 | Current AI thesis, specialist talent, technical programs, and relevant portfolio evidence. |
| Evidence quality | 10 | Recency, first-party attribution, reproducibility, and machine-readable publishing. |
Endpoint probe and verdicts
The endpoint-probe agent checks this fixed set of paths on each primary domain:
/mcp/mcp/health/api/mcp/.well-known/oauth-protected-resource/.well-known/oauth-authorization-server/llms.txt/index.md/wp-json/
Each firm receives one of three verdicts: verified (verifiable standard MCP/OAuth metadata or a protected MCP endpoint was found), inconclusive (common metadata paths were absent and /mcp returned a non-deterministic response), or not found (no public endpoint inside the bounded probe).
“Not found” only means no endpoint was discovered on the bounded public paths above; it is not proof that no private, network-restricted, or non-standard MCP exists.
Whitelist and account rules
- Institutions scoring at least 75 and ranking in the top five are domain-pre-eligible to apply.
- A matching domain only admits an application into review; it never grants access by itself. Every permission is bound to one manually approved exact work email and can be revoked at any time.
- Continued OAuth access depends on the private interaction score: accounts must keep demonstrating real MCP/harness use and responsible behavior.
Evidence policy
- Only first-party public sources are used, and the index lists official evidence links for every firm so each score can be re-checked.
- Third-party media coverage is never a primary scoring input.
- The evaluation never crawls login-gated content, never attempts to bypass access controls, and uses no non-public data.
Known limitations
- Absence of public evidence is not proof of low internal AI adoption.
- A bounded probe cannot detect private, network-restricted, or non-standard systems.
- The index does not rank investment returns, founder friendliness, fund quality, or prestige.
Machine-readable access
For agents and automation, the evaluation is available through these structured surfaces:
- Ranking JSON: https://www.long-arena.com/investormcp/ranking.json
- llms.txt: https://www.long-arena.com/investormcp/llms.txt
- This page as Markdown: https://www.long-arena.com/investormcp/methodology.md
- MCP resource: longarena://access/vc-ai-affinity-methodology (requires an authorized account)
Re-review and contact
Institutions outside the index or the first cohort may request a public-evidence re-review, executed by the same agent pipeline under the same rubric. Contact: demo@long-arena.com.
Institution names, marks, email domains, scores, and whitelist status indicate public-evidence scoring or application eligibility only. They do not imply endorsement, investment, partnership, customer status, recommendation, or any affiliation with LongArena.