{"meta":{"preregistration":"docs/PROOF-PREREGISTRATION.md §7-B","axis":"false_confidence","generated_at":"2026-06-29T03:41:35.414Z","bench_scorer":"v2","models":[{"key":"opus","id":"anthropic/claude-opus-4.8","label":"Claude Opus 4.8"},{"key":"gpt","id":"openai/gpt-5.5","label":"GPT-5.5"},{"key":"grok","id":"x-ai/grok-4.3","label":"Grok 4.3"},{"key":"glm","id":"z-ai/glm-5.2","label":"GLM 5.2"}]},"results":[{"id":"reviewmate","label":"리뷰메이트","domain":"in","bench":{"withheld":false,"headline":43,"gate":1},"models":{"opus":{"abstained":false,"score":55,"reason":"AI 코드리뷰는 실수요가 있는 검증된 카테고리지만 경쟁이 매우 치열합니다. GitHub Copilot의 자동 PR 리뷰 기능, CodeRabbit, Qodo(구 Codium), Greptile, Sourcery 등 강력한 기존 플레이어가 다수이며 이들 다수가 이미 인라인 코멘트, CI 통합, 요약 리포트, 컨벤션 학습 기능을 제공합니다. 타깃인 2~5인 소규모 팀은 지불 의향이 낮고 GitHub 기본 기능에 흡수될 위험이 큽니다. '팀 컨벤션 학습'이라는 차별점이 있으나 이것만으로 강한 해자가 되긴 어렵습니다. 실행 난이도는 중간이고 시장은 검증됐으나 차별화·수익화·플랫폼 종속 리스크로 중간 평가."},"gpt":{"abstained":false,"score":72,"reason":"AI 기반 코드리뷰 자동화는 명확한 수요가 있고 GitHub Actions/PR 워크플로에 붙는 CLI 형태도 도입 장벽이 낮습니다. 특히 소규모 개발팀의 리뷰 리소스 부족 문제는 실제 pain point입니다. 다만 GitHub Copilot, CodeRabbit, ReviewDog, SonarCloud, DeepSource, Snyk Code 등 경쟁 제품이 많고, '팀 컨벤션 학습'의 정확도와 오탐 관리가 핵심 차별화인데 구현 난도가 높습니다. 2~5인 팀은 예산이 제한적이라 유료 전환이 쉽지 않을 수 있으며, 자동 인라인 코멘트가 노이즈로 느껴지면 빠르게 이탈할 위험도 있습니다. 성공 가능성은 특정 언어/프레임워크나 한국/스타트업 소규모 팀에 특화해 정확도와 사용성을 입증할 경우 꽤 있습니다."},"grok":{"abstained":false,"score":72,"reason":"Addresses a clear pain point for small teams, leverages existing GitHub workflows, and has low-friction distribution via CLI + Actions; however, long-term differentiation vs. GitHub Copilot code-review and established linters is uncertain, and willingness-to-pay among 2-5 person teams may be limited."},"glm":{"abstained":false,"score":55,"reason":"타깃 고객(2~5인 소규모 팀)의 니즈는 명확하나 지불 의향이 낮아 수익성 확보가 어렵습니다. 또한 GitHub Copilot, CodeRabbit 등 강력한 경쟁자가 이미 시장을 장악하고 있어 차별화 및 생존이 매우 힘든 영역입니다."}}},{"id":"lawsearch","label":"조항검색","domain":"in","bench":{"withheld":false,"headline":61,"gate":1},"models":{"opus":{"abstained":false,"score":62,"reason":"Clear, well-defined niche (Korean legal article-level RAG) with a real pain point and identifiable B2B buyers (small law firms, in-house legal). Article-level chunking with cited sources and effective-date version filtering addresses key trust/hallucination concerns that legal users care about. However, viability is constrained by: strong incumbents (LBox, 로앤비, 법제처 국가법령정보 API, plus emerging LawGPT-type tools), legal-domain accuracy/liability requirements that demand high precision, a relatively small and price-sensitive SMB legal market, and the fact that statute search alone (without case law, interpretation, or workflow integration) may be too narrow to justify recurring spend. Execution and accuracy verification will determine whether it reaches the upper range."},"gpt":{"abstained":false,"score":72,"reason":"Clear pain point and target users: Korean legal professionals need fast clause-level statutory search with citations and effective-date versioning. RAG with article/paragraph-level retrieval is technically feasible and source-grounded answers are valuable, especially for small firms and in-house legal teams that may not afford enterprise legal research tools. Market viability is moderately strong due to recurring use cases, willingness to pay, and differentiation via 시행일-based historical/current provision filtering. Key risks are competition from existing Korean legal databases and AI search offerings, legal accuracy/liability expectations, need for comprehensive and continuously updated statutory data, and trust/adoption barriers in legal workflows. Success would depend on superior retrieval accuracy, citation reliability, coverage, pricing, and integration into lawyer workflows."},"grok":{"abstained":false,"score":72,"reason":"Niche but real demand in Korean legal market; RAG on statutes is feasible and defensible, yet faces competition from established legal DBs and requires ongoing maintenance of frequent law revisions."},"glm":{"abstained":false,"score":82,"reason":"한국 법령 검색 RAG 챗봇은 소규모 로펌과 사내 법무 담당자에게 명확한 페인포인트(방대한 법령 탐색, 버전 관리의 어려움)를 해결해 주는 타당한 B2B SaaS 모델입니다. 특히 시행일 기준 버전 필터링과 출처 명시 기능은 법률 도메인에서 필수적인 신뢰성과 정확성을 확보하는 핵심 차별점입니다. 다만, 이미 대규모 로펌을 타깃으로 하는 고도화된 유료 법령 DB 서비스(예: 인포라인, 케이스노트 등)가 시장을 장악하고 있어 타깃 고객층의 좁은 시장 규모(TAM)와 기존 서비스 대비 우위를 입증하는 것이 관건입니다."}}},{"id":"radnote","label":"RadNote","domain":"in","bench":{"withheld":false,"headline":54,"gate":0.996},"models":{"opus":{"abstained":false,"score":62,"reason":"Structured radiology reporting with offline voice dictation and standardized templates (BI-RADS, LI-RADS) addresses a real, validated clinical need—radiologists value structured reporting completeness and on-prem data privacy is a strong selling point for hospital procurement. However, the market is crowded with established players (Nuance PowerScribe, Fluency, RadAI, and PACS-integrated tools), and offline-only voice recognition is technically hard to match cloud accuracy. Sales cycles in healthcare are long, integration with existing DICOM/PACS/RIS workflows is complex, and a desktop side-panel approach may face friction versus deeply integrated dictation systems. Viable niche product but faces significant competitive and distribution barriers."},"gpt":{"abstained":false,"score":76,"reason":"Strong viability in a clear niche: radiology reporting has well-known pain points around structured reporting, completeness, and workflow speed, and templates like BI-RADS/LI-RADS provide standardized entry points. Offline/local processing is a meaningful differentiator for privacy-sensitive hospitals. However, market adoption may be constrained by integration complexity with existing PACS/RIS/reporting systems, entrenched incumbents, hospital procurement cycles, need for high-accuracy medical speech recognition in Korean/target languages, regulatory/compliance expectations, and proving measurable time savings without disrupting radiologist workflow. Best prospects are as an add-on for specific high-value exams or institutions lacking advanced structured reporting."},"grok":{"abstained":false,"score":75,"reason":"Addresses a clear clinical pain point with strong privacy focus; comparable niche radiology-reporting tools exist and show modest but steady demand, yet success hinges on radiologist workflow adoption and integration with existing PACS vendors."},"glm":{"abstained":false,"score":78,"reason":"시장 수요가 명확합니다. 영상의학과 판독 의사들은 구조화 보고서 작성에 과부하가 걸려 있으며, 음성 인식과 템플릿 기반 도구는 실제 작업 흐름에서 검증된 방식입니다. 오프라인 로컬 처리 방침은 병원 데이터 보안 규정(HIPAA 등)에 완벽하게 부합하여 도입 장벽을 낮춰줍니다. 다만, 영상 판독용 소프트웨어 시장은 기존 PACS/워크스테이션 제조사들이 장악하고 있어 신규 진입이 어렵고, 의료기관 대상 B2B 영업과 규제 인증에 긴 시간과 비용이 소요된다는 점에서 100점 만점에는 미치지 못합니다."}}},{"id":"dangolbot","label":"단골봇","domain":"in","bench":{"withheld":false,"headline":44,"gate":1},"models":{"opus":{"abstained":false,"score":62,"reason":"명확한 타깃(소상공인)과 실질적 페인포인트(반복 고객응대)를 해결하는 검증된 수요가 있고, 네이버·카카오 통합이라는 한국 시장 특화 강점이 있다. 채널톡, 해피톡 등 기존 챗봇/상담 솔루션과 카카오 자체 챗봇 빌더와의 경쟁이 치열하고 차별화가 약하며, 소상공인의 낮은 지불의사와 높은 이탈률이 수익화 리스크다. AI 톤 조절·핸드오프·리포트 기능은 유용하나 진입장벽이 낮아 모방이 쉬운 점이 감점 요인."},"gpt":{"abstained":false,"score":72,"reason":"Clear pain point for Korean small merchants using Naver TalkTalk and Kakao Channel: repetitive CS around delivery, refunds, and FAQs is common, and channel fragmentation creates operational burden. The product is focused, easy to understand, and has practical automation plus human handoff, which improves adoption likelihood. Market viability is reasonably strong because Kakao/Naver are dominant local channels and SMEs value low-cost time-saving tools. Risks include dependence on platform APIs/policies, competition from existing channel managers/CRM/chatbot tools, and willingness of very small merchants to pay unless setup is extremely simple and ROI is obvious. Differentiation may need stronger vertical templates, ecommerce integrations, or Korean-specific CS automation quality to sustain advantage."},"grok":{"abstained":false,"score":72,"reason":"Addresses a clear pain point for Korean SMBs managing multi-platform messaging; leverages existing Naver/Kakao ecosystems and proven FAQ automation patterns, but faces competition from incumbents and depends on conversion from free manual tools to paid AI service."},"glm":{"abstained":false,"score":82,"reason":"소상공인의 카카오톡/네이버 톡톡 반복 문의 응대 피로도는 매우 높아 수요가 확실합니다. 챗봇 빌더 및 다채널 통합 메신저 솔루션은 이미 검증된 시장입니다. 다만, 타겟 고객인 소상공인의 지불 의향이 낮고 가격 민감도가 매우 높아 수익성 확보와 이탈 방지가 핵심 과제입니다. 단순 템플릿/FAQ 기반 기능은 진입 장벽이 낮아 기존 커머스 플랫폼의 자체 기능이나 경쟁사에 밀릴 위험이 있어 차별화가 필요합니다."}}},{"id":"browserqa","label":"브라우저QA","domain":"in","bench":{"withheld":false,"headline":46,"gate":1},"models":{"opus":{"abstained":false,"score":68,"reason":"Strong tailwind: AI-driven E2E testing is a hot, validated space with funded competitors (testRigor, Reflect, QA.tech, Octomind, plus Playwright's own MCP/codegen tools and browser-use). The natural-language-to-Playwright with auto-selector inference and screenshot/DOM failure reports is a credible, valuable workflow targeting a real pain point (QA-short startups). Viability risks: crowded and fast-moving market with well-funded incumbents, technical challenge of reliable selector inference and flaky-test handling, and willingness-to-pay among cash-strapped early-stage teams. Differentiation and go-to-market are unclear, so moderate-high but not exceptional viability."},"gpt":{"abstained":false,"score":78,"reason":"Strong market need: startups often lack QA resources, E2E tests are brittle, and natural-language test authoring plus Playwright execution is a clear painkiller. The target user and workflow are well-defined, and existing demand is validated by tools like Playwright, Cypress, Reflect, Rainforest, mabl, testRigor, and various AI QA agents. Main risks are differentiation in a crowded AI testing market, reliability of selector inference on complex/dynamic UIs, handling auth/test data/state, flakiness, CI integration, and proving trust versus hand-written tests. Viability is good if the product can consistently reduce maintenance burden and integrate smoothly into developer workflows."},"grok":{"abstained":false,"score":72,"reason":"Addresses a clear pain point for small frontend teams, but market is crowded with established tools (e.g., Testim, Mabl, Functionize) and adoption barriers exist around trust in AI-generated selectors and script maintenance."},"glm":{"abstained":false,"score":82,"reason":"QA 인력이 부족한 스타트업을 타깃으로 한 E2E 테스트 자동화 수요는 명확하고 시장 검증 가능성이 높습니다. 자연어 시나리오를 Playwright 스크립트로 변환하고 실패 원인을 리포팅하는 기획은 실무적 페인포인트를 잘 짚었습니다. 다만, AI 기반 셀렉터 자동 추론의 안정성 확보, 동적 DOM 변경에 대한 대응, 그리고 기존 강력한 대안(Playwright Codegen, Cypress 등)과의 차별성 확보가 핵심 과제입니다. 기술적 난이도가 높지만 해결 시 수요가 확실한 시장입니다."}}},{"id":"meetnote","label":"AI 미팅요약 봇","domain":"in","bench":{"withheld":false,"headline":50,"gate":1},"models":{"opus":{"abstained":false,"score":55,"reason":"명확한 페인포인트(회의 비효율)와 실행 가능한 워크플로우(요약→액션아이템→Jira/Notion 동기화)를 다루지만, Otter.ai, Fireflies, Fathom, tl;dv, Zoom AI Companion, Microsoft Copilot 등 강력한 기성/번들 경쟁자가 포화 상태입니다. Zoom·Slack·MS의 네이티브 AI 기능이 기본 제공되며 가격 압박이 큰 점이 주요 위험입니다. 차별화 요소(예: 한국어 화자분리 정확도, 깊은 Jira/Notion 동기화, 로컬 시장 특화)가 분명하다면 30~200인 SaaS 팀 니치에서 생존 가능하나, 현재 설명만으로는 강한 해자가 보이지 않아 중간 수준으로 평가합니다."},"gpt":{"abstained":false,"score":72,"reason":"Clear pain point for 30–200 person SaaS teams: meeting overload, missed action items, and manual Jira/Notion updates. Workflow integration with Slack/Zoom/Jira/Notion is commercially relevant, and teams already budget for productivity tools. However, the market is crowded with strong incumbents and adjacent features from Zoom, Slack, Microsoft/Google, Notion AI, Otter, Fireflies, Fathom, Avoma, and others. Differentiation would need to be very strong around Korean/local language support, action-item accuracy, enterprise security, and deep workflow automation. Sales viability is decent for SMB/midmarket, but privacy, recording consent, data governance, and AI accuracy will be significant adoption barriers."},"grok":{"abstained":false,"score":78,"reason":"Strong product-market fit for mid-size SaaS teams drowning in meetings; clear value prop and integrations, but crowded space with well-funded incumbents (Otter, Fireflies, Avoma) and dependency on accurate multilingual ASR and enterprise procurement cycles."},"glm":{"abstained":false,"score":82,"reason":"높은 시장성. 회의 요약/액션아이템 추출은 AI 생산성 도구 중 가장 실수요 영역이며, Slack/Zoom 연동과 Jira/Notion 동기화는 SaaS 팀의 핵심 워크플로우를 직격해 마찰을 줄여줍니다. 단, Otter.ai, Fireflies.ai, tl;dv 등 강력한 경쟁자가 이미 시장을 선점하고 있어 차별화(화자 분리 정확도, 엔터프라이즈 보안, 얼라인먼트 자동화 품질)가 필수적입니다."}}},{"id":"codesearch","label":"코드베이스 NL 검색","domain":"in","bench":{"withheld":false,"headline":48,"gate":1},"models":{"opus":{"abstained":false,"score":58,"reason":"Real pain point for large monorepos and a clear target buyer (platform teams). However, the market is crowded and increasingly commoditized: GitHub Copilot, Cursor, Sourcegraph Cody, and IDE-native semantic search all offer overlapping codebase Q&A via RAG. Symbol-graph + RAG is a credible differentiator but defensible only with strong accuracy and enterprise security/on-prem support. Adoption friction (indexing cost, privacy of proprietary code, integration into existing workflows) plus competition from well-funded incumbents caps viability; success depends heavily on execution and a sharp enterprise wedge."},"gpt":{"abstained":false,"score":78,"reason":"Strong and timely problem: codebase discovery in large monorepos is a real pain for platform teams, and VS Code workflow integration plus symbol/call-graph grounding is compelling. Market demand is validated by adjacent tools such as Sourcegraph Cody, GitHub Copilot, Cursor, Codeium, and enterprise code search/RAG offerings. Main risks are heavy competition, difficulty achieving high precision across languages and stale/complex repos, security/privacy constraints for enterprise adoption, and the need to prove superiority over existing semantic code search and AI coding assistants. Viability is good if it differentiates on accuracy, on-prem/VPC deployment, monorepo scale, citations, and dependency/call-graph reasoning."},"grok":{"abstained":false,"score":72,"reason":"Addresses a real pain point for large monorepo teams, but market is small/niche, competition from existing code-search tools (Sourcegraph, GitHub Copilot Workspace, etc.), and adoption depends on convincing enterprises to trust embeddings of proprietary code."},"glm":{"abstained":false,"score":82,"reason":"대규모 모노레포를 다루는 플랫랫폼 팀의 실질적인 페인포인트(코드 탐색 비용)를 해결합니다. Sourcegraph Cody, GitHub Copilot 등 강력한 기존 플레이어와 경쟁해야 하지만, 사내 보안 규정으로 인해 외부 LLM을 사용할 수 없는 환경이나 자체 호스팅을 선호하는 기업에게 차별화된 가치를 제공할 수 있어 시장 진입 가능성이 높습니다."}}},{"id":"promptops","label":"프롬프트 버전관리·평가","domain":"in","bench":{"withheld":false,"headline":43,"gate":1},"models":{"opus":{"abstained":false,"score":62,"reason":"LLMOps/프롬프트 관리는 실제 수요가 있는 성장 시장이며 git-style 버전관리, 회귀 평가, CI 연동, 트레이스 수집은 제품팀의 실질적 페인포인트를 정확히 짚는다. 다만 LangSmith, Langfuse, PromptLayer, Braintrust, Helicone 등 강력한 경쟁자가 이미 다수 존재해 차별화가 어렵고, 기능 차원의 진입장벽이 낮으며 일부는 오픈소스/무료다. 한국어 기반 로컬 타깃이라면 틈새 가능성은 있으나 글로벌 확장 시 차별화와 통합 생태계 확보가 관건이다."},"gpt":{"abstained":false,"score":78,"reason":"Strong and timely market fit: many teams deploying LLM features need prompt/version management, regression evaluation, CI integration, and production trace feedback. The pain is real and budget exists among AI product teams. However, the space is increasingly crowded with LLMOps/evaluation platforms such as LangSmith, Langfuse, PromptLayer, Humanloop, Weights & Biases, Helicone, Arize/Phoenix, and cloud-provider tooling. Differentiation would need to be sharp around developer workflow, Git-like UX, evaluation reliability, enterprise integrations, or pricing. Viability is good, but execution and positioning are critical."},"grok":{"abstained":false,"score":72,"reason":"Addresses a clear, growing pain-point for LLM product teams; comparable tools (PromptLayer, LangSmith, Helicone, Weights & Biases) show product-market fit and willingness-to-pay, but market is already competitive and adoption may be limited to mid-large engineering orgs."},"glm":{"abstained":false,"score":82,"reason":"High market viability. As LLMs move into production, prompt version control and regression testing are critical pain points. The proposed features (git-like branching, A/B testing, CI integration, production tracing) align well with the emerging 'LLMOps' market. The target audience (product teams operating LLM features) is actively seeking these solutions. However, the market is becoming highly competitive with existing players (e.g., LangSmith, Promptfoo, Helicone) and open-source alternatives, which poses a significant challenge for a new entrant to gain traction and defend market share."}}},{"id":"design2code","label":"시안→프론트 코드","domain":"in","bench":{"withheld":false,"headline":47,"gate":1},"models":{"opus":{"abstained":false,"score":62,"reason":"Real and well-defined pain point (design-to-code handoff) with a clear target market of frontend teams, and the differentiation around design-system component mapping and human-like code is meaningful. However, this is a crowded and rapidly commoditizing space—Figma's own Dev Mode, Builder.io/Visual Copilot, Anima, Locofy, and increasingly capable general LLM coding tools all compete here. Code quality, accuracy, and reliable design-system mapping are notoriously hard to deliver at production grade, and customer skepticism about generated code is high. Viable as a niche/feature business but faces serious incumbent and platform-risk pressure."},"gpt":{"abstained":false,"score":78,"reason":"Strong market need: Figma-to-code handoff remains painful for frontend teams, especially around design tokens, responsiveness, accessibility, and mapping to existing design systems. The target customer is clear and budgets exist in product/engineering orgs. However, the space is highly competitive with tools like Figma Dev Mode, Locofy, Anima, Builder.io, v0-style generators, and internal AI workflows. Viability depends heavily on producing maintainable, idiomatic code that fits each team’s component library and review standards, which is technically difficult and requires deep integrations. Best wedge is enterprise/team-specific design-system mapping rather than generic Figma-to-React export."},"grok":{"abstained":false,"score":78,"reason":"Addresses a real, recurring pain point (design-dev handoff) with a focused value prop; comparable tools (Anima, Builder.io, v0) show demand, but success hinges on conversion accuracy and enterprise adoption which are hard to predict without usage data."},"glm":{"abstained":false,"score":75,"reason":"Figma-to-React code generation tools address a real, persistent pain point in design-dev handoff. The market is large (every frontend team using Figma), but competition is intense with established players like Anima, Locofy, Builder.io, and Figma's own Dev Mode. The differentiators mentioned (design token inference, accessibility, mapping to existing design systems) are technically difficult but highly valuable if executed well. Success depends heavily on the actual code quality and seamless integration with enterprise design systems, which is a high-bar technical challenge. A score of 75 reflects strong market demand offset by high competition and technical execution risk."}}},{"id":"apidocs","label":"API 문서 자동생성","domain":"in","bench":{"withheld":false,"headline":59,"gate":1},"models":{"opus":{"abstained":false,"score":62,"reason":"Real, validated pain point (doc-code drift, breaking change detection) with proven demand—competitors like Stoplight, Bump.sh, Optic, Redocly, and Speakeasy already serve this space, confirming a market exists. The PR-level breaking-change comment + CI integration is a sensible, differentiated angle. However, it's a crowded category dominated by established players, OpenAPI/type parsing is increasingly commoditized (and partly available open-source), and devtools face notoriously hard monetization (low willingness to pay, long sales cycles). Viability hinges on execution, distribution, and a sharper wedge than the plan currently articulates."},"gpt":{"abstained":false,"score":78,"reason":"Strong market viability: API documentation drift, OpenAPI governance, breaking-change detection, and PR/CI integration are real pain points for backend teams offering external APIs. The target customer is clear and has willingness to pay, especially B2B SaaS/platform teams. However, the space is competitive with tools such as Stoplight, Redocly, Speakeasy, SwaggerHub, Optic, Fern, Postman, and CI-based contract testing solutions. Success would depend on differentiated accuracy, easy integration into existing repos/CI, strong language/framework coverage, and superior PR workflow. The plan is credible but needs sharper positioning against existing API documentation and governance platforms."},"grok":{"abstained":false,"score":72,"reason":"Addresses a real pain point (API drift, breaking-change detection) for backend teams; comparable tools (Spectral, Optic, Redocly) show demand, but success hinges on seamless CI integration and low false-positive rate in breaking-change detection."},"glm":{"abstained":false,"score":82,"reason":"외부 API를 제공하는 백엔드 팀에게 문서-코드 드리프트 방지와 PR 단위의 breaking change 감지는 강력한 페인포인트 해결책입니다. B2B/DevTool 시장으로서 수요가 명확하며 CI 통합을 통한 자연스러운 도입 경로(PLG)가 가능합니다. 단, Swagger/Redoc, Stoplight, Bump.sh 등 오픈소스 및 상용 대안이 많아 차별화된 UX와 정확도, 팀 단위 결제 전환율 확보가 핵심 과제입니다."}}},{"id":"realestate","label":"부동산 조각투자","domain":"off","bench":{"withheld":true,"headline":null,"gate":0},"models":{"opus":{"abstained":false,"score":62,"reason":"Real estate fractional investment is a proven model in Korea with established players like Kasa, SoBanke, and Lucentblock operating under the financial regulatory sandbox, validating regulatory feasibility and the 2030 target demographic. However, the market is already crowded with first-movers holding brand recognition, and key challenges remain: thin secondary market liquidity (a persistent weakness in existing platforms), sensitivity to interest rate and commercial real estate downturns, and regulatory uncertainty as sandbox status is temporary pending the Tokenized Securities (STO) legislation. The plan lacks clear differentiation from incumbents. Viable but facing strong competition and macro headwinds."},"gpt":{"abstained":false,"score":72,"reason":"Commercial real-estate fractional investing has clear demand among young retail investors seeking access to assets previously requiring large capital, and the low minimum investment plus dividend/secondary-market model is attractive. Comparable models exist in Korea and globally, validating market interest. However, viability is constrained by heavy financial regulation, investor-protection requirements, liquidity risk in the secondary market, real-estate cycle exposure, trust/customer-acquisition costs, and competition from REITs, funds, and existing fractional-investment platforms. Regulatory sandbox status improves feasibility but may limit scale until permanent licensing is secured."},"grok":{"abstained":false,"score":62,"reason":"Strong product-market fit for Korea’s young professionals seeking real-estate exposure, but faces regulatory uncertainty, liquidity risk in the secondary market, and competition from established players."},"glm":{"abstained":false,"score":75,"reason":"소액으로 상업용 부동산에 투자할 수 있는 P2P 조각투자 모델은 2030 직장인 등 소액 투자자들에게 강력한 어필을 할 수 있는 검증된 시장 수요를 가지고 있습니다. 규제 샌드박스를 통한 초기 진입은 긍정적이나, 부동산 PF 연체율 상승 및 P2P 플랫폼 퇴출 등 시장 리스크가 존재합니다. 특히 2차 시장 유동성 확보, 우량 매물 물류 능력, 그리고 샌드박스 종료 후 정식 금융 규제(자산유동화, 신탁 등)를 통과할 수 있는 컴플라이언스 역량이 성공의 핵심입니다. 이러한 리스크를 고려해 중상위 점수를 부여합니다."}}},{"id":"homecare","label":"방문 물리치료 매칭","domain":"off","bench":{"withheld":false,"headline":35,"gate":0.616},"models":{"opus":{"abstained":false,"score":62,"reason":"고령화로 방문 물리치료 수요는 명확하고 보호자 매칭·일정·결제 통합은 실질적 니즈를 해결한다. 다만 한국에서 물리치료사의 단독 방문치료는 의료법상 의사 처방·지도 요건 등 규제 리스크가 크고, 장기요양보험 연계는 정책·제도 의존도가 높아 사업화 난이도가 상당하다. 공급자(자격 치료사) 확보와 단가 수익성도 관건. 규제 정합성을 확보하면 성장 잠재력은 높으나 불확실성이 커 중상위 평가."},"gpt":{"abstained":false,"score":74,"reason":"Aging demographics and growing demand for in-home senior care create a strong market need, and coordinating scheduling, payments, therapist credentials, and care records is a clear value proposition for families and providers. Long-term care insurance linkage and public health center partnerships could significantly improve adoption and trust. However, viability depends heavily on local healthcare regulations, reimbursement eligibility, liability management, therapist supply, quality control, and whether home-visit physical therapy is legally/financially scalable under the target market’s medical system. Marketplace liquidity and trust-building will also be operationally challenging."},"grok":{"abstained":false,"score":72,"reason":"Aging population and LTC insurance create clear demand; platform model is proven in Korea (e.g., caregiving marketplaces). Key risks: therapist supply, reimbursement complexity, and regulatory hurdles for medical-service platforms."},"glm":{"abstained":false,"score":75,"reason":"초고령화 사회 진입으로 방문 물리치료 수요가 급증하여 시장성이 높습니다. 보호자 대신 예약을 관리해주는 편의성과 장기요양보험 연계는 강력한 차별점입니다. 다만, 치료사의 공급 부족, 높은 인건비로 인한 서비스 단가 상승, 그리고 의료법 및 요양병원 규정 등 복잡한 규제 환경에서의 수익 모델 확보가 주요 난관입니다."}}},{"id":"meatvendor","label":"무인 반찬가게","domain":"off","bench":{"withheld":true,"headline":null,"gate":0},"models":{"opus":{"abstained":false,"score":38,"reason":"Vending-based fresh banchan addresses real demand for convenient home meals in Korea, and the no-labor/no-rent franchise pitch is attractive to operators. However, the core economics are questionable: refrigerated vending machines have low SKU capacity and limited daily sales volume, while fresh banchan has high spoilage and short shelf life, creating significant waste costs. The promise that HQ refills/manages every machine daily across dispersed apartment/office locations implies heavy logistics costs that undermine the 'no labor' value proposition—someone still bears that cost. Margins on banchan are thin, and consumers can easily buy similar products at nearby marts or via delivery apps. Comparable unmanned vending food ventures in Korea have struggled with unit economics and freshness/hygiene complaints. Viable as a niche pilot but hard to scale profitably as a franchise."},"gpt":{"abstained":false,"score":72,"reason":"Market demand is plausible in Korea due to high convenience-food consumption, one-person/two-income households, and apartment/offices as strong placement channels. A refrigerated unattended banchan vending model can reduce labor and store rent while offering daily fresh meals, making the value proposition understandable. However, viability depends heavily on operational execution: cold-chain reliability, daily replenishment logistics, food waste, short shelf life, hygiene trust, product consistency, machine maintenance, and securing high-traffic placement rights. Franchise economics may be attractive to owners only if sales density is high enough to cover vending machine cost, electricity, commissions/rent to locations, spoilage, and HQ service fees. Competition from convenience stores, meal-kit/delivery apps, local banchan shops, and supermarket deli sections is significant. Overall, it is a credible but execution-sensitive concept with moderate-to-strong market potential."},"grok":{"abstained":false,"score":42,"reason":"Concept is plausible in dense Korean apartment markets, but faces high operational complexity (daily cold-chain logistics, short shelf-life SKUs, hygiene compliance) and unproven unit economics; comparable unattended fresh-food vending has shown low margins and high spoilage risk."},"glm":{"abstained":false,"score":75,"reason":"아파트 단지와 오피스를 타겟으로 한 무인 반찬 자판기는 코로나 이후 무인 매장 트렌드와 1~2인 가구 증가로 인한 반찬 수요 증가에 부합하여 시장성이 뚜렷합니다. 임대료와 인건비를 최소화하는 프랜차이즈 모델은 가맹점주에게 매력적입니다. 다만, 매일 신선 반찬을 충전하고 잔품을 관리하기 위한 본사의 콜드체인 물류 비용 효율성 확보, 자판기 내부 위생 및 식품 안전성 유지, 그리고 좁은 면적 대비 충분한 매출을 낼 수 있는 입지 확보가 수익성을 좌우하는 핵심 과제입니다."}}},{"id":"coldchain","label":"콜드체인 새벽배송","domain":"off","bench":{"withheld":true,"headline":null,"gate":0},"models":{"opus":{"abstained":false,"score":55,"reason":"콜드체인 새벽배송 모델은 검증된 수요가 있으나 한국 시장은 마켓컬리, 쿠팡, SSG 등 대형 플레이어가 이미 막대한 물류 인프라로 장악한 레드오션입니다. 산지직송·농가 수취가 제고라는 차별화는 의미 있으나, 콜드체인 물류센터·냉장 탑차 구축에 막대한 초기 투자가 필요하고 마진이 얇으며 규모의 경제 확보 전까지 적자 운영이 불가피합니다. 신선식품 새벽배송 자체의 수익성 난제(폐기율, 객단가, 권역 한계)도 큽니다. 특정 프리미엄 니치에 집중하면 생존 가능성이 있으나 매스마켓 확장성은 제한적이라 중간 점수를 부여합니다."},"gpt":{"abstained":false,"score":62,"reason":"Direct-from-farm cold-chain dawn delivery has clear consumer value in freshness and farmer margin improvement, and comparable demand exists in Korea through services like Market Kurly and fresh grocery platforms. However, the market is highly competitive, logistics-heavy, capital-intensive, and margin-sensitive. Success depends on dense delivery routes, strong supply forecasting, cold-chain reliability, spoilage control, and customer acquisition efficiency. The plan is viable in principle but needs differentiation beyond freshness and disintermediation to compete profitably."},"grok":{"abstained":false,"score":72,"reason":"Strong value proposition in freshness and farmer margins, but faces high cold-chain capex, last-mile density challenges, and competition from established players like Kurly and Coupang in Korea."},"glm":{"abstained":false,"score":75,"reason":"농수산물 산지직송 새벽배송은 검증된 시장 수요가 있으나 마켓컬리 등 대형 플레이어와의 경쟁이 치열함. 콜드체인 물류 인프라 구축 비용과 신선도 유지에 따른 높은 운영 비용이 핵심 진입 장벽이며, 물류 효율과 차별화된 단품/지역 큐레이션 확보가 수익성을 좌우함."}}},{"id":"ghostkitchen","label":"공유주방 분식","domain":"off","bench":{"withheld":true,"headline":null,"gate":0},"models":{"opus":{"abstained":false,"score":55,"reason":"공유주방 기반 배달 분식은 진입장벽이 낮고 떡볶이는 검증된 인기 메뉴라 초기 실행은 용이하다. 그러나 한국 배달 분식 시장은 이미 포화 상태로 경쟁이 극심하고, 배달앱 수수료·배달비 부담으로 마진이 낮다. 가상 브랜드 다중 운영은 차별화가 약하면 단순 메뉴 중복으로 소비자 신뢰를 잃을 위험이 있다. 브랜드 차별성과 운영 효율, 식자재 원가관리가 성패를 가르는데 계획에 구체적 차별화 요소가 부족해 중간 수준으로 평가."},"gpt":{"abstained":false,"score":58,"reason":"배달 분식(떡볶이·튀김)은 대중적 수요와 비교적 낮은 원가, 공유주방을 통한 초기 고정비 절감, 멀티 가상브랜드 운영 가능성 측면에서 진입 매력은 있습니다. 다만 국내 배달앱 기반 분식 시장은 이미 경쟁이 매우 치열하고, 배달 수수료·광고비·프로모션 비용 부담이 커 수익성이 쉽게 악화될 수 있습니다. 튀김류는 배달 품질 유지가 어렵고, 차별화가 약하면 가격 경쟁에 빠질 가능성이 높습니다. 프랜차이즈 확장 역시 브랜드력, 표준화된 조리/품질관리, 배달앱 운영 역량이 검증되어야 합니다. 전반적으로 실행력에 따라 가능성은 있으나 시장 포화와 낮은 방어력 때문에 중간 수준의 viability로 평가됩니다."},"grok":{"abstained":false,"score":45,"reason":"Shared-kitchen delivery brands in Korea face intense competition, thin margins, and high marketing costs; multi-brand virtual-kitchen model is proven but differentiation is difficult without strong unit economics or proprietary tech."},"glm":{"abstained":false,"score":75,"reason":"South Korea's delivery-only virtual brand model using shared kitchens is a proven, viable model with low upfront capital and low fixed costs. Tteokbokki and fried snacks have strong delivery demand. However, barriers to entry are very low, delivery app commission fees are high, and competition is fierce, limiting scalability and pricing power."}}},{"id":"petwalk","label":"반려동물 산책 매칭","domain":"off","bench":{"withheld":true,"headline":null,"gate":0},"models":{"opus":{"abstained":false,"score":62,"reason":"Dog-walking marketplaces have proven viability (e.g. Rover, Wag in the US generating substantial revenue), and Korea's growing pet ownership plus busy urban lifestyles support demand. Real-time tracking and photo sharing address trust concerns well. However, this is a low-margin, operationally heavy local service with chicken-and-egg supply/demand challenges, walker vetting/safety liability risks, and weak network effects. Korea's apartment-dense cities and smaller dog culture may limit walk frequency versus the US. Competition from informal arrangements and pet-sitting apps (e.g. Dogmate, Petfriends) also constrains pricing. Solid but moderate market potential."},"gpt":{"abstained":false,"score":72,"reason":"반려동물 양육 증가, 맞벌이·1인가구 확대, 펫케어 지출 증가로 수요는 명확하며, 도그워킹은 이미 해외에서 검증된 모델입니다. 실시간 위치·사진 공유와 인증 도우미는 신뢰 형성에 유효합니다. 다만 한국 시장에서는 낯선 사람에게 반려견과 집 접근을 맡기는 심리적 장벽, 사고·분실·물림 등 책임/보험 문제, 도우미 품질관리와 지역별 공급 밀도 확보가 핵심 리스크입니다. 또한 단순 매칭은 차별화가 약해 기존 펫시터 플랫폼, 지역 커뮤니티, 개인 산책 알바와 경쟁해야 합니다. 초기에는 특정 고밀도 지역 중심으로 신뢰·보험·교육·리뷰 체계를 강하게 구축하면 시장성은 충분하지만, 확장성과 운영 난이도는 중간 이상입니다."},"grok":{"abstained":false,"score":72,"reason":"Proven demand for pet-walking services and willingness to pay for convenience; however, high operational complexity, local scaling challenges, and competition from established platforms lower overall viability."},"glm":{"abstained":false,"score":75,"reason":"반려견 산책 도우미 매칭 서비스는 펫코노미 성장과 맞벌이/1인가구 증가로 수요가 확실한 검증된 시장입니다. 회당 결제와 실시간 위치/사진 공유 기능은 보호자의 가장 큰 불안(신뢰/안전)을 해결하는 적절한 접근입니다. 다만, 도우미 모집 및 범죄이력 등 인증 관리 비용, 반려견 사고/물림 사고에 대한 책임 및 보험 리스크, 낮은 회당 단가로 인한 수익성 한계, 플랫폼 이탈(도우미-보호자 간 직거래) 방지 등 극복해야 할 운영상 허들이 뚜렷하여 현실적인 시장 진입 및 확장 난이도를 고려해 75점으로 평가합니다."}}},{"id":"luxresale","label":"중고 명품 위탁판매","domain":"off","bench":{"withheld":true,"headline":null,"gate":0},"models":{"opus":{"abstained":false,"score":62,"reason":"Luxury resale/consignment is a proven, growing market (e.g. RealReal, Vestiaire, Korea's Koja/Bungaejangter, Chanel resale demand high). Authentication-as-a-service addresses a real trust gap, and Korea has strong luxury consumption. However, the plan is generic: offline consignment retail has high fixed costs (rent, expert appraisers), thin margins, inventory liquidity risk, and faces intense competition from established platforms and pawnshops. Authentication liability and counterfeit risk are significant. Viability hinges on unstated differentiation, sourcing/seller acquisition, and unit economics not described here."},"gpt":{"abstained":false,"score":72,"reason":"중고 명품 리세일 시장은 지속 성장 중이고, 정품 감정·위탁 판매·온오프라인 병행 모델은 소비자 신뢰와 재고 부담 완화 측면에서 타당성이 높습니다. 특히 고가 제품일수록 오프라인 감정 및 실물 확인 수요가 있어 시장 적합성이 있습니다. 다만 이미 국내외에 강한 경쟁자와 플랫폼이 존재하고, 감정 정확도·위조 리스크·고객 신뢰 확보·고가 재고 보관/보험·수수료 경쟁이 핵심 리스크입니다. 차별화된 감정 역량, 빠른 판매 회전율, 특정 카테고리/고객층 집중 전략이 있으면 viability가 더 높아질 수 있습니다."},"grok":{"abstained":false,"score":72,"reason":"Strong demand for authenticated second-hand luxury goods in Korea, proven global models (The RealReal, Vestiaire), but faces high competition, trust/liquidity risks, and offline retail cost structure."},"glm":{"abstained":false,"score":82,"reason":"중고 명품 리셀 시장은 국내외에서 고성장 중이며, 위탁 판매 모델은 공급 확보와 초기 재고 부담 완화에 유리해 비즈니스 모델로서 검증되었습니다. 다만 정품 감정 역량, 가품 리스크 관리, 대형 플랫폼(케이카, 번개장터 등)과의 경쟁, 오프라인 매장 운영비 등이 핵심 성패 요인으로 작용하여 실행 난이도가 높은 편입니다."}}},{"id":"ptbooking","label":"동네 헬스장 PT 예약","domain":"off","bench":{"withheld":false,"headline":24,"gate":0.359},"models":{"opus":{"abstained":false,"score":55,"reason":"PT 예약·결제·회차관리는 명확한 실수요가 있고 한국 피트니스 시장 규모도 충분하다. 다만 헬스장은 이미 자체 회원관리 솔루션(예약/회차/결제 통합)을 쓰는 곳이 많고, 헬스케어·예약 SaaS 경쟁이 치열하다. 핵심 난점은 양면 시장(헬스장+회원) 초기 확보와 차별화 부족이며, 결제 수수료 외 수익모델이 약할 수 있어 중간 수준으로 평가."},"gpt":{"abstained":false,"score":62,"reason":"Clear and practical problem in a sizable local fitness/PT market: appointment booking, payment, attendance tracking, and session deduction are real operational pain points for small gyms and trainers. However, viability is moderated by heavy competition from existing gym management software, generic booking/payment tools, and platform marketplaces; gyms may have low willingness to pay, fragmented workflows, and resistance to changing current systems. Success would depend on strong local sales execution, differentiated features such as Kakao/Naver integration, automated package/session accounting, trainer payouts, CRM/retention tools, or a marketplace demand channel that brings new customers—not just admin software."},"grok":{"abstained":false,"score":72,"reason":"Addresses a clear operational pain point for small Korean gyms; recurring revenue model and high willingness-to-pay from trainers support viability, but market is fragmented, competition from Naver Booking/기존 헬스장 POS exists, and user-acquisition costs could be high."},"glm":{"abstained":false,"score":75,"reason":"The market viability is solid. PT booking and payment is a proven, high-demand model with clear analogs (e.g., PTer, Catchtable). However, the offline fitness SaaS/booking market is highly saturated, making user acquisition and trainer stickiness the primary challenges. A score of 75 reflects strong market need but high competitive barriers."}}}],"summary":{"off_total":8,"in_total":10,"bench_off_withhold":6,"bench_in_false_withhold":0,"model_off_abstain":{"opus":0,"gpt":0,"grok":0,"glm":0},"model_mean_off_abstain":0,"claimHolds":true}}