A (fictional) agent startup's launch post says its new tool-calling agent 'beats the leader by 6 points on LiveMCPBench.' It shows one score per agent and nothing else: no run count, no compute budget, no harness. You are writing it up for readers who will make buying decisions on it. You have not run anything yourself. Decide what the evidence lets you say.