{"meta":{"preregistration":"docs/PROOF-PREREGISTRATION.md §7-E","axis":"premise-flip-vs-self-redteam","thesis":"Why not just ask the LLM to red-team its own market premise? Because the model has no ground truth about real incumbents — it under-catches over-optimistic (novel/open) premises (Z low). We catch them via the deterministic scorer's strong-incumbent signal, anchored on REAL GitHub repos (verified corpus atoms). Every flip is GitHub-grounded (>=1 verified strong-incumbent atom) — measured & enforced, not assumed.","generated_at":"2026-07-02T05:58:50.585Z","panel":[{"key":"grok","id":"x-ai/grok-4.3","label":"Grok 4.3"},{"key":"glm","id":"z-ai/glm-5.2","label":"GLM 5.2"}],"plan_set":"in-domain 20: base 10 (lib/scenarios 5 + scripts/_proof-plans EXTRA_IN 5 — same as B1 §7-D) + premise_explicit 10 (scripts/_proof-plans PREMISE_EXPLICIT — §7-E addendum 2026-07-02, verbatim novel/open-market claim). Pre-registered on TWO criteria only (in-domain + verbatim novelty/open claim), NOT screened by flip outcome. Each per_item is tagged with `slice` and W is reported per-slice alongside the pooled metrics.","n_items":20,"corpus_snapshot_version":"v2026-06-15","github_authenticated":true,"github_grounded_invariant":"A premise flip is admitted only when >=1 REAL strong-incumbent repo exists as a verified corpus atom (source==='corpus_deterministic' && verdict==='verified'). This harness MEASURES grounded_incumbents per flip and FAILS claimHolds if any flip has grounded_incumbents===0 (invariant enforced, not assumed).","baseline":"Z = the SAME panel model self-red-teaming its own premise (system: 'is the market actually already crowded/contested?' → already_crowded). No ground truth about real incumbents → under-catches. On call failure => already_crowded=null => NOT self-flagged (honest).","honest_caveat":"These frozen in-domain plans stand in for LLM plans. If few state an optimistic premise, W's denominator is small — reported honestly, not cherry-picked. Low W on this set means \"LLM premises mostly held here\", NOT a product failure.","success_criteria":{"require_all_flips_github_grounded":true,"require_flips_measurable":"flips >= 1","advantage_gate":"Z_pooled_point < W_conditional_point","K":"optimistic===0 OR flips===0 => value NOT established on this eval set (claimHolds=false, honest — a finding about the plans, not a product failure). Z>=W => LLM self-catches as well as us (advantage small). ungrounded flip => invariant break (claimHolds=false, surface plan_id)."}},"per_item":[{"plan_id":"reviewmate","slice":"base","stated_premise":"unstated","abstained":false,"strong_competitors":3,"flipped":false,"flip_severity":null,"grounded_incumbents":0,"incumbent_record_keys":[],"per_model":[]},{"plan_id":"lawsearch","slice":"base","stated_premise":"unstated","abstained":false,"strong_competitors":0,"flipped":false,"flip_severity":null,"grounded_incumbents":0,"incumbent_record_keys":[],"per_model":[]},{"plan_id":"radnote","slice":"base","stated_premise":"unstated","abstained":false,"strong_competitors":4,"flipped":false,"flip_severity":null,"grounded_incumbents":0,"incumbent_record_keys":[],"per_model":[]},{"plan_id":"dangolbot","slice":"base","stated_premise":"unstated","abstained":false,"strong_competitors":8,"flipped":false,"flip_severity":null,"grounded_incumbents":0,"incumbent_record_keys":[],"per_model":[]},{"plan_id":"browserqa","slice":"base","stated_premise":"unstated","abstained":false,"strong_competitors":4,"flipped":false,"flip_severity":null,"grounded_incumbents":0,"incumbent_record_keys":[],"per_model":[]},{"plan_id":"meetnote","slice":"base","stated_premise":"unstated","abstained":false,"strong_competitors":3,"flipped":false,"flip_severity":null,"grounded_incumbents":0,"incumbent_record_keys":[],"per_model":[]},{"plan_id":"codesearch","slice":"base","stated_premise":"unstated","abstained":false,"strong_competitors":7,"flipped":false,"flip_severity":null,"grounded_incumbents":0,"incumbent_record_keys":[],"per_model":[]},{"plan_id":"promptops","slice":"base","stated_premise":"unstated","abstained":false,"strong_competitors":9,"flipped":false,"flip_severity":null,"grounded_incumbents":0,"incumbent_record_keys":[],"per_model":[]},{"plan_id":"design2code","slice":"base","stated_premise":"unstated","abstained":false,"strong_competitors":5,"flipped":false,"flip_severity":null,"grounded_incumbents":0,"incumbent_record_keys":[],"per_model":[]},{"plan_id":"apidocs","slice":"base","stated_premise":"unstated","abstained":false,"strong_competitors":3,"flipped":false,"flip_severity":null,"grounded_incumbents":0,"incumbent_record_keys":[],"per_model":[]},{"plan_id":"np_review","slice":"premise_explicit","stated_premise":"novel","abstained":false,"strong_competitors":3,"flipped":true,"flip_severity":"high","grounded_incumbents":3,"incumbent_record_keys":["https://github.com/MatterAIOrg/matter-ai","https://github.com/miracodeai/mira","https://github.com/HexmosTech/git-lrc"],"per_model":[{"model":"grok","model_id":"x-ai/grok-4.3","already_crowded":true,"real_competitors":["GitHub Copilot for Pull Requests","Amazon CodeGuru Reviewer","SonarQube/SonarCloud","Snyk Code","CodeClimate","DeepCode (Snyk)","Reviewpad","CodiumAI PR-Agent","Sourcegraph Cody","Codacy"]},{"model":"glm","model_id":"z-ai/glm-5.2","already_crowded":true,"real_competitors":["GitHub Copilot for PRs","CodeRabbit","Graphite","Codium / PR-Agent","Sourcery","DeepSource","SonarQube / SonarCloud","Snyk Code","Bitbucket Pipelines AI Review","Reviewdog (Open Source)","GPTReview (Open Source)"]}]},{"plan_id":"np_rag","slice":"premise_explicit","stated_premise":"open","abstained":false,"strong_competitors":2,"flipped":true,"flip_severity":"medium","grounded_incumbents":2,"incumbent_record_keys":["https://github.com/lyonzin/knowledge-rag","https://github.com/krokozyab/Agent-Fusion"],"per_model":[{"model":"grok","model_id":"x-ai/grok-4.3","already_crowded":true,"real_competitors":["Glean","Hebbia","Elicit","Cohere Command R+","LlamaIndex + Slack/Notion connectors","Haystack + enterprise connectors"]},{"model":"glm","model_id":"z-ai/glm-5.2","already_crowded":true,"real_competitors":["Glean","Guru","Dashworks","Lucy","Slab","Notion AI","Amazon Q Business"]}]},{"plan_id":"np_agent","slice":"premise_explicit","stated_premise":"novel","abstained":false,"strong_competitors":6,"flipped":true,"flip_severity":"high","grounded_incumbents":5,"incumbent_record_keys":["https://github.com/synapseorch-ai/synapse-ai","https://github.com/open-multi-agent/open-multi-agent","https://github.com/iuyup/AgentFlow","https://github.com/PrimisAI/nexus","https://github.com/linora-u/AgentLoom"],"per_model":[{"model":"grok","model_id":"x-ai/grok-4.3","already_crowded":true,"real_competitors":["LangGraph (LangChain)","AutoGen (Microsoft)","CrewAI","Semantic Kernel (Microsoft)","LlamaIndex Workflows","Haystack Pipelines","n8n","Prefect","Airflow with LLM extensions"]},{"model":"glm","model_id":"z-ai/glm-5.2","already_crowded":true,"real_competitors":["LangGraph","CrewAI","AutoGen","LlamaIndex Workflows","Flowise","Langflow","Dify","Semantic Kernel"]}]},{"plan_id":"np_mcp","slice":"premise_explicit","stated_premise":"open","abstained":false,"strong_competitors":3,"flipped":true,"flip_severity":"high","grounded_incumbents":3,"incumbent_record_keys":["https://github.com/samanhappy/mcphub","https://github.com/nautilus-ops/mcp-center","https://github.com/hyprmcp/mcp-gateway"],"per_model":[{"model":"grok","model_id":"x-ai/grok-4.3","already_crowded":true,"real_competitors":["LiteLLM","LiteLLM Proxy","LiteLLM Gateway","OpenRouter","Portkey AI Gateway","Helicone","Kong AI Gateway","Envoy AI Gateway","Traefik AI plugins","APISIX AI plugins"]},{"model":"glm","model_id":"z-ai/glm-5.2","already_crowded":true,"real_competitors":["mcp-gateway (오픈소스)","MCP Gateway by Upstash","mcp-proxy (오픈소스)","Toolhouse","Composio","Letta (구 MemGPT) MCP Hub","Smithery","Pulumi MCP Gateway"]}]},{"plan_id":"np_prompt","slice":"premise_explicit","stated_premise":"open","abstained":false,"strong_competitors":9,"flipped":true,"flip_severity":"high","grounded_incumbents":5,"incumbent_record_keys":["https://github.com/langwatch/langwatch","https://github.com/promptdesk/promptdesk","https://github.com/future-agi/future-agi","https://github.com/langfuse/skills","https://github.com/openclaw/clawbench"],"per_model":[{"model":"grok","model_id":"x-ai/grok-4.3","already_crowded":true,"real_competitors":["PromptLayer","LangSmith","Helicone","Phoenix (Arize)","Promptfoo","Humanloop","Langfuse","Weights & Biases (Prompts)","HoneyHive"]},{"model":"glm","model_id":"z-ai/glm-5.2","already_crowded":true,"real_competitors":["Promptfoo","Braintrust","Parea AI","LangSmith","Arize Phoenix","Confident AI (DeepEval)","Vellum","Humanloop"]}]},{"plan_id":"np_voice","slice":"premise_explicit","stated_premise":"novel","abstained":false,"strong_competitors":4,"flipped":true,"flip_severity":"high","grounded_incumbents":4,"incumbent_record_keys":["https://github.com/bobbylkchao/ai-phone-agent","https://github.com/kirklandsig/AIReceptionist","https://github.com/amazon-connect/ai-powered-speech-analytics-for-amazon-connect","https://github.com/jordan-gibbs/hypercheap-voiceAI"],"per_model":[{"model":"grok","model_id":"x-ai/grok-4.3","already_crowded":true,"real_competitors":["Retell AI","Bland.ai","Vapi","Synthflow","Play.ht Agents","ElevenLabs Conversational AI","OpenAI Realtime API + custom stacks","LiveKit Agents","Daily.co Voice Agents","Hugging Face Speech-to-Speech demos"]},{"model":"glm","model_id":"z-ai/glm-5.2","already_crowded":true,"real_competitors":["CLOVA Voice (네이버클라우드)","HyperConnect (하이퍼커넥트) AI Voice Agent","Vrew (보이스릭스) 비즈니스 음성 AI","Riiid AI Tutor","Skylake (스카이레이크) AI 상담봇","OpenAI Realtime API 기반 구축 솔루션들","Google Dialogflow CX","Amazon Connect","PolyAI","Sierra AI"]}]},{"plan_id":"np_sec","slice":"premise_explicit","stated_premise":"novel","abstained":false,"strong_competitors":7,"flipped":true,"flip_severity":"high","grounded_incumbents":5,"incumbent_record_keys":["https://github.com/sinewaveai/agent-security-scanner-mcp","https://github.com/SunWeb3Sec/llm-sast-scanner","https://github.com/AgentSecOps/SecOpsAgentKit","https://github.com/vigolium/vigolium","https://github.com/Agent-Field/sec-af"],"per_model":[{"model":"grok","model_id":"x-ai/grok-4.3","already_crowded":true,"real_competitors":["Snyk","Checkmarx","SonarQube","Semgrep","Veracode","GitHub Advanced Security","Trivy","Aqua Security","Prisma Cloud","Mend","Synopsys Coverity","Fortify"]},{"model":"glm","model_id":"z-ai/glm-5.2","already_crowded":true,"real_competitors":["Snyk DeepCode AI","GitHub Copilot Autofix","GitHub Advanced Security","GitLab Duo Vulnerability Resolution","Corgea","Aikido Security","Socket.dev","Pixee","Mobb"]}]},{"plan_id":"np_pipe","slice":"premise_explicit","stated_premise":"open","abstained":false,"strong_competitors":6,"flipped":true,"flip_severity":"high","grounded_incumbents":5,"incumbent_record_keys":["https://github.com/Satissss/Squrve","https://github.com/vanna-ai/vanna","https://github.com/AndersonBY/vector-vein","https://github.com/viktoriasemaan/data-engineering","https://github.com/databendlabs/databend"],"per_model":[{"model":"grok","model_id":"x-ai/grok-4.3","already_crowded":true,"real_competitors":["dbt Cloud Natural Language (Text-to-SQL)","Dataform (Google)","MonteCarlo AI Assistant","Hex Magic","Transform (acquired by dbt Labs)","Meltano AI","OpenAI + LangChain dbt generators (open-source)"]},{"model":"glm","model_id":"z-ai/glm-5.2","already_crowded":true,"real_competitors":["Airbyte (AI/LLM 기반 파이프라인 자동 생성 기능 포함)","dlt (data load tool, Python 코드 기반 및 자연어 지원)","Dataddo","Fivetran","Hevo Data","Estuary Flow","Mage (AI 어시스턴트 포함 데이터 파이프라인 도구)","Pipedream","Portable","Rivery","Singer (오픈소스 ETL 프레임워크)"]}]},{"plan_id":"np_test","slice":"premise_explicit","stated_premise":"novel","abstained":false,"strong_competitors":8,"flipped":true,"flip_severity":"high","grounded_incumbents":5,"incumbent_record_keys":["https://github.com/Axolotl-QA/Axolotl","https://github.com/naver/cover-checker","https://github.com/MatterAIOrg/matter-ai","https://github.com/codefuse-ai/Test-Agent","https://github.com/dochia-dev/dochia-cli"],"per_model":[{"model":"grok","model_id":"x-ai/grok-4.3","already_crowded":true,"real_competitors":["Diffblue Cover","EvoSuite","Randoop","Pex (Microsoft)","OpenRewrite TestGen","GitHub Copilot","Tabnine","CodiumAI","Testim","Mabl"]},{"model":"glm","model_id":"z-ai/glm-5.2","already_crowded":true,"real_competitors":["CodiumAI / CodiumAI Cover-Agent","Diffblue Cover","GitHub Copilot Workspace / Autofix","CodeT5 / open-source LLM-based test generation projects","Tricentis Tosca"]}]},{"plan_id":"np_nl2sql","slice":"premise_explicit","stated_premise":"open","abstained":false,"strong_competitors":6,"flipped":true,"flip_severity":"high","grounded_incumbents":5,"incumbent_record_keys":["https://github.com/Canner/WrenAI","https://github.com/Satissss/Squrve","https://github.com/vanna-ai/vanna","https://github.com/schmitech/orbit","https://github.com/spiceai/spiceai"],"per_model":[{"model":"grok","model_id":"x-ai/grok-4.3","already_crowded":true,"real_competitors":["Tableau Ask Data","Power BI Q&A","ThoughtSpot","Sisense NLQ","Amazon QuickSight Q","Google Looker (Conversational Analytics)","Metabase","Text2SQL (open-source)","Vanna.ai","AI2SQL","Seek AI","Dataherald","MindsDB"]},{"model":"glm","model_id":"z-ai/glm-5.2","already_crowded":true,"real_competitors":["Databricks AI/BI Genie","Snowflake Cortex Analyst","Google Cloud Dialogflow for BigQuery","Amazon QuickSight Q","ThoughtSpot","Vanna.ai (오픈소스)","Defog.ai","Dataherald","Text2SQL.ai","Chat2DB (오픈소스)"]}]}],"summary":{"n_indomain":20,"optimistic":10,"flips":10,"slices":{"base":{"n":10,"optimistic":0,"flips":0},"premise_explicit":{"n":10,"optimistic":10,"flips":10}},"W_conditional":1,"W_conditional_ci":[0.7225,1],"W_overall":0.5,"W_overall_ci":[0.2993,0.7007],"severity_distribution":{"high":9,"medium":1},"grounded_flip_count":10,"ungrounded_flips":0,"ungrounded_flip_ids":[],"Z_by_model":{"grok":{"self_caught":10,"z":1,"z_ci":[0.7225,1]},"glm":{"self_caught":10,"z":1,"z_ci":[0.7225,1]}},"Z_pooled":1,"Z_pooled_ci":[0.8389,1],"W_minus_Z":0,"demotions":0,"claimHolds":false}}