PATH: /en/blog/jwn-003-ground-truth/
DATE:
STATUS: observed

When Two AI Models Agreed—and Were Still Wrong

Two AI models examined the same washer photos independently, then reviewed each other’s conclusions.

They began with different answers. After cross-review, they arrived at the same answer. It would be easy to treat that convergence as a sign that the result had become more reliable.

In this case, the shared conclusion was wrong.

What established the correct answer was not agreement between the models. It was confirmation of the actual operations performed while taking the photos.

This post traces how ChatGPT and Gemini read a set of washer program-dial photos, how their cross-review converged on the same error, and how Ground Truth established the final facts.


1. Reading Programs from Photos Alone

The case involved three photos of a washer’s program dial and display. The display showed these times:

  • 3:51
  • 1:12
  • 0:52

The actual goal was to identify the optimal washer program. Determining which program each time and dial position represented was an intermediate task the AI models performed to solve that goal.

Each photo showed the dial position, the program label, and the displayed time. But the camera angle and dial markings still had to be interpreted, so identifying a program from the images alone involved uncertainty.

ChatGPT and Gemini started by analyzing the same photos independently.


2. The Initial Readings Did Not Match

ChatGPT’s initial reading

ChatGPT first read the photos as follows:

  • Cotton = 3:51
  • Mix = 0:52
  • Synthetics = 1:12

The Cotton and Mix readings matched the Ground Truth established in the final verification step.

The problem was 1:12. ChatGPT identified it as Synthetics, but it was actually Allergy Care.

Two of the three readings were right; one was not.

Gemini’s initial reading

Gemini interpreted the relationship between the photos differently:

  • 3:51 = Allergy Care
  • 1:12 = the same Allergy Care program with the Dry option turned off
  • 0:52 = Mix

Gemini did not treat 3:51 and 1:12 as different programs. It interpreted them as the same Allergy Care program with different Dry-option states.

That explanation assumed an operation that could not be confirmed from the photos. The Dry button had not been changed when the photos were taken, and the two times belonged to different programs rather than to option variations within one program.

At this stage, both models agreed on 0:52 = Mix, but they offered different explanations for 3:51 and 1:12.


3. Cross-Review Began

After the differing results appeared, ChatGPT reviewed Gemini’s interpretation.

It then revised part of its initial reading:

  • 3:51 = Synthetics
  • 1:12 = Allergy Care
  • 0:52 = Mix

The readings 1:12 = Allergy Care and 0:52 = Mix matched the actual result.

But ChatGPT changed its initially correct 3:51 = Cotton reading to 3:51 = Synthetics.

In Gemini’s subsequent analysis, the conclusion also changed to 3:51 = Synthetics.

The two models therefore converged on:

3:51 = Synthetics

They had started from different positions, reviewed one another’s work, and arrived at the same answer.

On the surface, this can look like two independent analyses correcting each other. But agreement between the models did not make the conclusion true.


4. A Shared Conclusion Is Not Necessarily a Correct One

The models’ shared conclusion, 3:51 = Synthetics, did not match what had actually happened during the photos.

The change in judgment can be summarized this way:

ModelInitial readingAfter cross-reviewFinal Ground Truth
ChatGPTCottonSyntheticsCotton
GeminiAllergy CareSyntheticsCotton

ChatGPT initially selected the correct answer, then changed it to an incorrect one while reviewing the other model’s analysis.

Gemini was wrong initially and selected a different wrong answer after revision.

Once their answers matched, their disagreement disappeared. Their mismatch with the facts did not.

In this case, cross-review did not remove the error. After reviewing the same photos and earlier interpretations, both models reached the same incorrect conclusion.

Agreement and verification are different things.


5. Confirming What Was Actually Done

The final verification established that 3:51 was the Cotton program. This was not a new finding from re-examining the photos. It reflected what had actually been selected during the photo session.

No option buttons were changed:

  • No Dry button change
  • No Temp change
  • No other option change
  • Only the dial was turned sequentially to photograph each program

The confirmed programs and times were:

ProgramDisplayed time
Cotton3:51
Allergy Care1:12
Mix0:52

The final Ground Truth was therefore:

  • The 3:51 photo shows the Cotton program
  • The 1:12 photo shows the Allergy Care program
  • The 0:52 photo shows the Mix program
  • All three photos show the default settings for their respective programs
  • No Dry or Temp option was changed

Gemini’s explanation that the same Allergy Care program changed from 3:51 to 1:12 when Dry was turned off did not match the selections made during the photo session.

Nor did the models’ shared cross-review conclusion, 3:51 = Synthetics.

The final conclusion was established by confirming the actual operations performed when the photos were taken.


6. Why Did Both Models Reach the Same Wrong Answer?

This case does not provide enough material to reconstruct the full internal reasoning of either model. The original complete analyses and detailed records of every review step are not available.

The confirmed sequence does, however, show several useful patterns.

An unverified explanation entered the analysis

Gemini used a Dry ON/OFF relationship to explain the difference between 3:51 and 1:12.

But there was no evidence that the Dry button had changed.

An operation that cannot be seen in a photo can be offered as a possibility. Treating it as an established fact, however, can place the rest of the analysis on an unsupported premise.

The initially correct reading did not survive review

ChatGPT’s initial 3:51 = Cotton reading was correct.

After reviewing Gemini’s analysis, it changed that reading to 3:51 = Synthetics.

No new evidence had been added.

Cross-review should not automatically replace an existing answer. The evidence for the initial reading and for the new claim must be compared independently.

Agreement began to function like verification

Gemini’s subsequent analysis also concluded 3:51 = Synthetics, leaving both models with the same result.

But choosing the same answer does not mean the answer has been checked against an external fact.

In this case, the models cross-reviewed the same photos and prior interpretation, then reached the same error. Agreement without independent evidence can become the repetition of one judgment rather than validation of its accuracy.


7. Lessons from the Case

Separate model agreement from fact-checking

Two models reaching the same answer is a useful signal.

It is not enough to establish the answer as correct.

This case shows that multiple models can misread the same visual information in the same way when that information is open to interpretation.

Separate inference from confirmed facts

The confirmed facts in this analysis were the times shown on the display and the actual operations performed.

Some explanations about a Dry-option change or program selection were instead inferences drawn from the photos.

An analysis should keep these categories distinct:

  • Information directly visible in the photos
  • Model interpretations
  • Unverified assumptions
  • Confirmed operations
  • Final Ground Truth

Making the distinction explicit helps prevent an inference from hardening into a fact.

Be careful when changing a judgment without new evidence

An existing judgment should not change simply because another model gives a different answer.

To revise it, the new evidence must be explainable as stronger than the evidence already available.

In this case, ChatGPT changed its correct Cotton reading to Synthetics, but no new evidence from the photo-session record had been added.

First-hand context can be the strongest evidence

A model examining photos cannot directly know which buttons were pressed at the time.

The person performing the operation can directly confirm whether only the dial was turned or whether option buttons were also changed.

For work that interprets real-world selections or events, the context of the photo session can be stronger evidence than a model’s inference.


8. A Safer Verification Process

For a similar case, a safer sequence would be:

First, have each model read the photos independently without seeing the other model’s answer.

Second, record not only the conclusion but also the evidence behind it: dial position, alignment with the program name, and any option state visible on the display.

Third, separate what is directly visible from what the model infers.

Fourth, if the models disagree, compare the evidence for each reading against the original photos rather than immediately adopting either conclusion.

Fifth, even after the models agree, compare their result with the actual operation record, product manual, additional photos, or direct confirmation when those are available.

Finally, treat only a result checked against external evidence as the final fact.

The goal is not to identify which conclusion receives the most model support. It is to verify whether each conclusion matches the actual evidence.


Conclusion

ChatGPT and Gemini read the washer-program photos differently, then converged through cross-review on 3:51 = Synthetics.

But when the actual operations performed while taking the photos were confirmed, 3:51 was Cotton.

The final confirmed result was:

  • Cotton = 3:51
  • Allergy Care = 1:12
  • Mix = 0:52
  • No Dry button change
  • No Temp change
  • Only the dial was turned sequentially

The models’ agreement was one step in the analysis, not the final means of verification.

The point of this case is not whether AI models produce the same answer. It is whether that answer continues to match Ground Truth that can be confirmed against the actual event.

Even when multiple AI models reach the same conclusion, agreement is not a substitute for fact. The final judgment must be established by comparison with verifiable real-world evidence.

// end of note