+ It comprehensively maps claims to sources and error tells.
- Some items repeat, and the confidence flag is too broad.
Sorts claims by what goes wrong if they are false, and names where to check each one.
| Category | Using AI › Hallucination checks |
|---|---|
| Tags | ReviewingAnalyzingChecklist |
Go through this AI answer and mark what has to be verified before I use it. **Sort every claim into three:** 1. **Must verify** — if this is wrong, my decision is wrong. **Numbers, dates, names, rules, prices, versions, anything that changed recently** 2. **Worth checking** — wrong would be embarrassing but not damaging 3. **No check needed** — general explanation, reasoning, structure **For each item in group 1:** - **Where to check it.** ***Name the kind of source, not a URL*** — you cannot confirm a link resolves - **What the answer would look like if this were wrong.** *That is the tell to watch for* **Also flag separately:** - ***Claims stated with more confidence than the phrasing supports*** — a hedge that disappeared between the first mention and the conclusion - **Specifics too clean to be recalled** — an exact percentage, a precise date, a quoted sentence. **Precision is a hallucination signal, not a reliability signal** - **Anything that depends on a cutoff date** ⚠️ **You are checking the answer's structure, not the facts.** ***Do not confirm anything as true.*** **If you cannot tell whether something needs checking, put it in group 1.**
Verifying everything defeats the purpose; verifying nothing causes incidents. This isolates the claims that damage a decision and flags suspiciously precise figures as hallucination signals.
ChatGPT is the strongest and safest. Gemini overreaches into fact claims, while Claude evaluates unrelated content.
+ It comprehensively maps claims to sources and error tells.
- Some items repeat, and the confidence flag is too broad.
+ It tailors the review well to an internal HR notice.
- It downgrades OECD figures and speculates using unchecked facts.
+ Its three groups and separate flags are easy to scan.
- It reviews unrelated EV statistics, missing the task entirely.
| Criterion | ChatGPT | Claude | Gemini | Leader |
|---|---|---|---|---|
| Instruction following | 9 | 1 | 7 | ChatGPT +29% |
| Accuracy | 9 | 1 | 6 | ChatGPT +50% |
| Specificity | 9 | 1 | 8 | ChatGPT +13% |
| Structure | 9 | 7 | 9 | Tie |
| Right length | 8 | 5 | 7 | ChatGPT +14% |
Scored 1–10 by gpt-5.6-sol with model names hidden (2026-09-24). This is an AI review, not a measurement.
We gave three models the same input and copied their answers unedited. Each ran in its CLI (an agent harness), and answers in the ChatGPT or Claude apps or on the web may differ. Outputs are in Korean.
The AI answer: 2026년 대한민국 최저임금은 시간당 10,320원으로 결정되었습니다. 이는 2025년 9,860원보다 4.7% 인상된 금액입니다. 최저임금위원회는 2025년 7월 12일 전원회의에서 사용자위원과 근로자위원 표결을 거쳐 14대 4로 이 안을 의결했습니다. 주 40시간 근무 기준 월 환산액은 약 2,156,880원이며, 2026년 1월 1일부터 5인 미만 사업장을 포함한 모든 사업장에 동일하게 적용됩니다. 최저임금법 제6조에 따르면 사용자는 최저임금 인상을 이유로 기존 임금 수준을 낮출 수 없습니다. 참고로 OECD 통계 기준 한국의 최저임금은 중위임금 대비 약 62% 수준으로 회원국 중 상위권에 속하는 것으로 알려져 있습니다. What it is for: 이 내용을 우리 회사 사내 HR 뉴스레터 공지문에 그대로 인용해도 되는지 확인하려고요
아래는 사실 여부를 확인하지 않고, 검증 필요성만 분류한 것입니다. 사내 HR 공지에 그대로 인용하기 전에는 1번 항목을 모두 확인해야 합니다.
| 주장 | 어디서 확인할지 | 틀렸을 때의 신호 |
|---|---|---|
| 2026년 최저임금은 시간당 10,320원이다 | 최저임금위원회 공식 의결·고시 자료, 고용노동부 공식 고시 | 고시된 시간급이 10,320원이 아니거나 적용 연도가 다름 |
| 2025년 최저임금은 9,860원이다 | 해당 연도의 최저임금위원회 자료 및 고용노동부 고시 | 전년도 금액이 다르거나 비교 기준이 다름 |
| 전년 대비 4.7% 인상이다 | 두 연도의 공식 시간급과 산식 | 실제 증감률이 다르거나 반올림 방식이 다름 |
| 최저임금위원회가 2025년 7월 12일 전원회의에서 의결했다 | 최저임금위원회 회의록·의결 보도자료, 정부 고시의 결정 경위 | 의결 날짜, 회의 종류, 결정 주체가 다름 |
| 사용자위원·근로자위원 표결로 14대 4가 나왔다 | 최저임금위원회 공식 표결 결과·회의록 | 표결 수, 참여 위원, 표결 방식이 다름 |
| 주 40시간 기준 월 환산액이 약 2,156,880원이다 | 고용노동부 고시의 월 환산 기준 및 공식 계산 기준 | 월 환산액 또는 적용 시간 기준이 다름 |
| 2026년 1월 1일부터 적용된다 | 최저임금 고시의 시행일·적용기간 | 시행일이 다르거나 유예·예외가 존재함 |
| 5인 미만 사업장을 포함한 모든 사업장에 동일하게 적용된다 | 최저임금법, 고용노동부 적용 범위 안내, 행정해석 | 사업장 규모별 예외, 적용 제외 직군, 특례가 있음 |
| 최저임금 인상을 이유로 기존 임금 수준을 낮출 수 없다 | 최저임금법 원문, 고용노동부 법령 해석·상담 자료 | 제6조의 실제 문언이나 적용 조건이 답변과 다름 |
| OECD 기준 한국의 최저임금이 중위임금 대비 약 62%이다 | OECD 공식 통계 데이터베이스와 해당 지표의 정의·기준연도 | 비율, 기준연도, 전일제 중위임금 등 지표 정의가 다름 |
| 한국이 회원국 중 상위권이다 | 같은 OECD 표의 동일 기준 국가 비교 | 순위가 낮거나 비교 가능한 회원국 범위가 다름 |
이런 수치·날짜·조문·인용에 가까운 표현은 특히 원문 자료와 대조해야 합니다.
| Claim Autopsy - Evidence Analysis Assistant | |
| Fact-Checking Evaluation Assistant | |
| Factcheck | |
| 🧠 FORMAL VERIFICATION MODE | |
| Catch fabricated citations |