AI Incident Database: Tougher Turing Test Exposes Chatbots’ Stupidity (migrated to Issue)
The 2016 Winograd Schema Challenge highlighted how even the most successful AI systems entered into the Challenge were only successful 3% more often than random chance. This incident has been downgraded to an issue as it does not meet current ingestion criteria.
AI Incident Database · Incident 21Operator lens (medium): inspect D2 (output validation), D6 (autonomy oversight).
Event summary
The 2016 Winograd Schema Challenge highlighted how even the most successful AI systems entered into the Challenge were only successful 3% more often than random chance. This incident has been downgraded to an issue as it does not meet current ingestion criteria.
Linked entities
- researchers
organisation | 70%
Related graph edges
| Edge | Type | Confidence |
|---|---|---|
| ent-aiid-researchers to ent-psf-d2 | maps to | 60% |
| ent-aiid-researchers to ent-psf-d2 | maps to | 60% |
| ent-aiid-researchers to ent-psf-d2 | maps to | 60% |
| ent-aiid-researchers to ent-psf-d6 | maps to | 60% |
| ent-aiid-researchers to ent-psf-d6 | maps to | 60% |
| ent-aiid-researchers to ent-psf-d6 | maps to | 60% |