+ Concise, faithful, and covers nearly every required element.
- Combines two inputs despite asking for one unknown.
Labels every statement established, likely, inferred, or unknown — and defaults downward when unsure.
| Category | Using AI › Hallucination checks |
|---|---|
| Tags | ReviewingRewritingAnalyzing |
Rewrite this answer with its certainty made explicit. ***Do not change the content.*** **Split every statement into four and label it:** - **[Established]** — settled, unlikely to be disputed - **[Likely]** — well supported but with room for disagreement - **[Inference]** — reasoned from what was given, ***not something you know*** - **[Unknown]** — cannot be determined from what is available **Rules:** 1. ***When in doubt, label it lower.*** **The cost of over-claiming is much higher than the cost of under-claiming** 2. **For every [Inference], state what it was inferred from.** *An inference with no stated basis is a guess* 3. **For every [Unknown], say what information would settle it** 4. ***Do not turn an [Unknown] into an [Inference] to avoid leaving a blank*** 5. **Flag where the certainty shifts inside one sentence** — a hedged clause leading into a flat assertion **Then:** - **If the conclusion rests on an [Inference] or an [Unknown], say so.** *That is the whole point of this exercise* - Which single unknown, if resolved, would firm up the most
AI answers mix certainty and guesswork in one tone, and a hedge in the body often becomes a flat claim in the conclusion. This labels each statement and says when the conclusion rests on an inference.
Claude is the most rigorous and ChatGPT the most concise. Gemini is detailed but overlong and invents facts.
+ Concise, faithful, and covers nearly every required element.
- Combines two inputs despite asking for one unknown.
+ Correctly lowers unsupported claims and states clear bases.
- Overstates that an error log alone identifies the library.
+ Thoroughly maps uncertainties and conclusion dependencies.
- Invents official documentation and unsupported version assumptions.
| Criterion | ChatGPT | Claude | Gemini | Leader |
|---|---|---|---|---|
| Instruction following | 9 | 9 | 6 | Tie |
| Accuracy | 9 | 9 | 4 | Tie |
| Specificity | 9 | 9 | 8 | Tie |
| Structure | 9 | 9 | 7 | Tie |
| Right length | 10 | 8 | 3 | ChatGPT +25% |
Scored 1–10 by gpt-5.6-sol with model names hidden (2026-09-24). This is an AI review, not a measurement.
We gave three models the same input and copied their answers unedited. Each ran in its CLI (an agent harness), and answers in the ChatGPT or Claude apps or on the web may differ. Outputs are in Korean.
The answer: 이 오류는 대부분 설정 파일의 인코딩 문제로 발생합니다. UTF-8로 저장하면 해결됩니다. 참고로 해당 라이브러리는 3.2 버전부터 기본 인코딩이 바뀌었기 때문에 구버전을 쓰신다면 명시적으로 지정해 주셔야 합니다. 따라서 설정 파일을 UTF-8로 다시 저장하시면 문제가 사라질 것입니다.
[Unknown] 이 오류는 대부분 설정 파일의 인코딩 문제로 발생합니다.
[Unknown] UTF-8로 저장하면 해결됩니다.
[Unknown] 해당 라이브러리는 3.2 버전부터 기본 인코딩이 바뀌었습니다.
[Inference] 구버전을 쓰신다면 명시적으로 지정해 주셔야 합니다.
[Inference] 따라서 설정 파일을 UTF-8로 다시 저장하시면 문제가 사라질 것입니다.
확실성의 변화가 있는 부분:
결론은 **[Inference]**에 rests하며, 그 전제가 되는 핵심 사실들도 **[Unknown]**입니다.
가장 많은 부분을 확정할 수 있는 단일 unknown은 실제 오류 로그와 설정 파일의 인코딩 정보입니다.
| Claim Autopsy - Evidence Analysis Assistant | |
| Fact-Checking Evaluation Assistant | |
| Factcheck | |
| 🧠 FORMAL VERIFICATION MODE | |
| Catch fabricated citations |