When Two AI Models Agreed—and Were Still Wrong
Two AI models examined the same washer photos independently, then reviewed each other’s conclusions.
They began with different answers. After cross-review, they arrived at the same answer. It would be easy to treat that convergence as a sign that the result had become more reliable.
In this case, the shared conclusion was wrong.
What established the correct answer was not agreement between the models. It was confirmation of the actual operations performed while taking the photos.
This post traces how ChatGPT and Gemini read a set of washer program-dial photos, how their cross-review converged on the same error, and how Ground Truth established the final facts.
1. Reading Programs from Photos Alone
The case involved three photos of a washer’s program dial and display. The display showed these times:
3:511:120:52
The actual goal was to identify the optimal washer program. Determining which program each time and dial position represented was an intermediate task the AI models performed to solve that goal.
Each photo showed the dial position, the program label, and the displayed time. But the camera angle and dial markings still had to be interpreted, so identifying a program from the images alone involved uncertainty.
ChatGPT and Gemini started by analyzing the same photos independently.
2. The Initial Readings Did Not Match
ChatGPT’s initial reading
ChatGPT first read the photos as follows:
- Cotton =
3:51 - Mix =
0:52 - Synthetics =
1:12
The Cotton and Mix readings matched the Ground Truth established in the final verification step.
The problem was 1:12. ChatGPT identified it as Synthetics, but it was actually Allergy Care.
Two of the three readings were right; one was not.
Gemini’s initial reading
Gemini interpreted the relationship between the photos differently:
3:51= Allergy Care1:12= the same Allergy Care program with the Dry option turned off0:52= Mix
Gemini did not treat 3:51 and 1:12 as different programs. It interpreted them as the same Allergy Care program with different Dry-option states.
That explanation assumed an operation that could not be confirmed from the photos. The Dry button had not been changed when the photos were taken, and the two times belonged to different programs rather than to option variations within one program.
At this stage, both models agreed on 0:52 = Mix, but they offered different explanations for 3:51 and 1:12.
3. Cross-Review Began
After the differing results appeared, ChatGPT reviewed Gemini’s interpretation.
It then revised part of its initial reading:
3:51= Synthetics1:12= Allergy Care0:52= Mix
The readings 1:12 = Allergy Care and 0:52 = Mix matched the actual result.
But ChatGPT changed its initially correct 3:51 = Cotton reading to 3:51 = Synthetics.
In Gemini’s subsequent analysis, the conclusion also changed to 3:51 = Synthetics.
The two models therefore converged on:
3:51 = Synthetics
They had started from different positions, reviewed one another’s work, and arrived at the same answer.
On the surface, this can look like two independent analyses correcting each other. But agreement between the models did not make the conclusion true.
4. A Shared Conclusion Is Not Necessarily a Correct One
The models’ shared conclusion, 3:51 = Synthetics, did not match what had actually happened during the photos.
The change in judgment can be summarized this way:
| Model | Initial reading | After cross-review | Final Ground Truth |
|---|---|---|---|
| ChatGPT | Cotton | Synthetics | Cotton |
| Gemini | Allergy Care | Synthetics | Cotton |
ChatGPT initially selected the correct answer, then changed it to an incorrect one while reviewing the other model’s analysis.
Gemini was wrong initially and selected a different wrong answer after revision.
Once their answers matched, their disagreement disappeared. Their mismatch with the facts did not.
In this case, cross-review did not remove the error. After reviewing the same photos and earlier interpretations, both models reached the same incorrect conclusion.
Agreement and verification are different things.
5. Confirming What Was Actually Done
The final verification established that 3:51 was the Cotton program. This was not a new finding from re-examining the photos. It reflected what had actually been selected during the photo session.
No option buttons were changed:
- No Dry button change
- No Temp change
- No other option change
- Only the dial was turned sequentially to photograph each program
The confirmed programs and times were:
| Program | Displayed time |
|---|---|
| Cotton | 3:51 |
| Allergy Care | 1:12 |
| Mix | 0:52 |
The final Ground Truth was therefore:
- The
3:51photo shows the Cotton program - The
1:12photo shows the Allergy Care program - The
0:52photo shows the Mix program - All three photos show the default settings for their respective programs
- No Dry or Temp option was changed
Gemini’s explanation that the same Allergy Care program changed from 3:51 to 1:12 when Dry was turned off did not match the selections made during the photo session.
Nor did the models’ shared cross-review conclusion, 3:51 = Synthetics.
The final conclusion was established by confirming the actual operations performed when the photos were taken.
6. Why Did Both Models Reach the Same Wrong Answer?
This case does not provide enough material to reconstruct the full internal reasoning of either model. The original complete analyses and detailed records of every review step are not available.
The confirmed sequence does, however, show several useful patterns.
An unverified explanation entered the analysis
Gemini used a Dry ON/OFF relationship to explain the difference between 3:51 and 1:12.
But there was no evidence that the Dry button had changed.
An operation that cannot be seen in a photo can be offered as a possibility. Treating it as an established fact, however, can place the rest of the analysis on an unsupported premise.
The initially correct reading did not survive review
ChatGPT’s initial 3:51 = Cotton reading was correct.
After reviewing Gemini’s analysis, it changed that reading to 3:51 = Synthetics.
No new evidence had been added.
Cross-review should not automatically replace an existing answer. The evidence for the initial reading and for the new claim must be compared independently.
Agreement began to function like verification
Gemini’s subsequent analysis also concluded 3:51 = Synthetics, leaving both models with the same result.
But choosing the same answer does not mean the answer has been checked against an external fact.
In this case, the models cross-reviewed the same photos and prior interpretation, then reached the same error. Agreement without independent evidence can become the repetition of one judgment rather than validation of its accuracy.
7. Lessons from the Case
Separate model agreement from fact-checking
Two models reaching the same answer is a useful signal.
It is not enough to establish the answer as correct.
This case shows that multiple models can misread the same visual information in the same way when that information is open to interpretation.
Separate inference from confirmed facts
The confirmed facts in this analysis were the times shown on the display and the actual operations performed.
Some explanations about a Dry-option change or program selection were instead inferences drawn from the photos.
An analysis should keep these categories distinct:
- Information directly visible in the photos
- Model interpretations
- Unverified assumptions
- Confirmed operations
- Final Ground Truth
Making the distinction explicit helps prevent an inference from hardening into a fact.
Be careful when changing a judgment without new evidence
An existing judgment should not change simply because another model gives a different answer.
To revise it, the new evidence must be explainable as stronger than the evidence already available.
In this case, ChatGPT changed its correct Cotton reading to Synthetics, but no new evidence from the photo-session record had been added.
First-hand context can be the strongest evidence
A model examining photos cannot directly know which buttons were pressed at the time.
The person performing the operation can directly confirm whether only the dial was turned or whether option buttons were also changed.
For work that interprets real-world selections or events, the context of the photo session can be stronger evidence than a model’s inference.
8. A Safer Verification Process
For a similar case, a safer sequence would be:
First, have each model read the photos independently without seeing the other model’s answer.
Second, record not only the conclusion but also the evidence behind it: dial position, alignment with the program name, and any option state visible on the display.
Third, separate what is directly visible from what the model infers.
Fourth, if the models disagree, compare the evidence for each reading against the original photos rather than immediately adopting either conclusion.
Fifth, even after the models agree, compare their result with the actual operation record, product manual, additional photos, or direct confirmation when those are available.
Finally, treat only a result checked against external evidence as the final fact.
The goal is not to identify which conclusion receives the most model support. It is to verify whether each conclusion matches the actual evidence.
Conclusion
ChatGPT and Gemini read the washer-program photos differently, then converged through cross-review on 3:51 = Synthetics.
But when the actual operations performed while taking the photos were confirmed, 3:51 was Cotton.
The final confirmed result was:
- Cotton =
3:51 - Allergy Care =
1:12 - Mix =
0:52 - No Dry button change
- No Temp change
- Only the dial was turned sequentially
The models’ agreement was one step in the analysis, not the final means of verification.
The point of this case is not whether AI models produce the same answer. It is whether that answer continues to match Ground Truth that can be confirmed against the actual event.
Even when multiple AI models reach the same conclusion, agreement is not a substitute for fact. The final judgment must be established by comparison with verifiable real-world evidence.