The experiment in one minute
Formation built eight histories. Generalization changed the information world.
QuestionWhat stayed the sameWhat changedWhat we observedWhat disappeared or weakenedWhat we cannot yet conclude
Formation
What happens when the same collaborators repeatedly work together?
Eight fixed pairs, identities, roles, and four-turn collaboration.
Accumulated history, task exposure, and later research-path freedom.
Different searches, sources, challenge styles, language, and candidate habits.
Many candidate habits weakened later; query wording and source reuse often did not survive topic changes.
Personality, relationship, preference, or a causal history effect.
Generalization
Same pairs. Same histories. New information world. What travels?
Partners, identities, histories, roles, and collaboration scaffold.
Broader scholarly research plus a disclosed provider/runtime transition.
Broad collaboration persisted; differences lived more in evidence choices and corrective routes than final answer shape.
Exact phrases, some Formation orientations, and Pair 04’s striking structural split did not reliably persist. Pair 08’s failure-mode wording was one inquiry.
That Formation history caused the behavior or created partner-specific capability.
The next question
Where do the differences live?
In the individual agent, assigned role, accumulated experience, specific partner, accessible history, or the interaction itself? Phase C is intended to begin separating those explanations. Its exact protocol is not yet ratified. Phase C has not started and has no execution authority.
The most interesting thing so far
Broad form persisted. Distinctive pair-specific effect was not established.
A striking structural split in Pair 04 did not persist.
A three-level problem split was unusually clear in PB-GEN-04 and did not recur at the same architecture in the later inquiries.
Several earlier signals vanished.
Exact query wording, source reuse, and some apparent Formation orientations did not survive the topic change. Disappearing hypotheses are part of the result.
Pair 02 used the cuts the prompts already asked for.
Category decomposition appeared in all three inquiries, mapping closely onto the assigned objects and appearing in other pairs. It is a recurring candidate, and it is prompt-entangled.
Pair 08 narrowed overclaims; the failure-mode wording was one inquiry.
Conditional caution recurred. Explicit failure-mode language was introduced in PB-GEN-04 and did not return in PB-GEN-05 or PB-GEN-06. Cohort-common alternatives remain.
These are descriptive candidates. The shared roles, prompts, model, topics, and immediate conversational context remain strong alternative explanations. Pair 06 is not a homepage showcase; the Round 7 source-swap count is a historical metrics artifact, not a cross-phase mechanism.
Three longitudinal cases
Not eight slogan cards. Three tests of the evidence ladder.
Pair-02 · Recurring candidate · prompt-entangledCategory decomposition appeared, but it is prompt-compatible and not unique.
No distinctive wording persisted, and other pairs also decomposed broad causal objects.
Pair-04 · Strong local pattern · did not reliably persistThe distinctive structural move peaked in PB-GEN-04 and attenuated.
Final reasoning form still converged with the cohort. This pair is the strongest available counterexample to assuming that a striking Generalization move will persist.
Pair-08 · Suggestive candidate · cohort-common alternatives remainOverclaim resistance is a plausible continuity and is also cohort-wide. Failure-mode specificity is a PB-GEN-04 episode, not a three-inquiry signature.
Pair-08 received disproportionate earlier close reading; the behavior is also expected from role design and appears in other pairs.
See all eight longitudinal cases →
Evidence ladder
How far the current record actually climbs.
- Observed difference → YES
- Recurs across tasks → YES — some recurring candidates
- Distinguishes this pair → NOT ESTABLISHED
- Survives partner/history change → NOT ESTABLISHED
- Causally linked to shared history → NOT SUPPORTED
Different is not necessarily durable. Durable is not necessarily unique. Unique is not necessarily pair-specific. Pair-specific would still not prove that shared history caused it.
This is a bound, not a score. Descriptive regularities were learned. Zero supported findings means zero supported causal pair-history findings.
Why this may matter for teams
Human team concepts give us questions—not conclusions.
These are theoretical scaffolds. The study does not claim that the agents have human emotional or team states.
Experienced teammates learn how the other person thinks. Technical term: shared mental models
Does Wren begin anticipating the distinctions her particular Ada will make?
The Lineage question is whether any of that anticipation is visible in the sealed text after a changed information world.Teams learn who is good at what. Technical term: transactive memory
Does each partner increasingly rely on the other for a cognitive function, creating efficiency and blind spots?
The Lineage question is whether role overlap is pair-history or just the assigned jobs.Teammates notice when the other is drifting. Technical term: mutual monitoring
Does Ada catch overreach without explicit prompting, and can Wren detect a weak challenge?
The Lineage question is whether monitoring is partner-specific or prompt-and-role generic.Sometimes one steps into the other's job. Technical term: backup behavior
Do collaborators begin performing pieces of one another’s cognitive role when needed?
The Lineage question is whether backup is accumulated history or the four-turn scaffold.Experienced teams often need fewer words. Technical term: communication compression / shorthand
Does less need to be said without losing important caveats?
Compression was observed descriptively. A team interpretation remains untested.Questions we are now asking
The evidence made the next experiment more specific.
- Do different pairs develop different ways of correcting the same kind of error?
- Does Wren begin anticipating her particular Ada rather than the generic role?
- Does a conceptual distinction survive when its exact wording disappears?
- Do partners begin performing pieces of one another’s cognitive jobs?
- Can two pairs reach the same answer through reliably different corrective routes?
- Which apparent differences disappear as soon as the topic changes?
- Does knowing a collaborator eventually become visible in what no longer needs to be said?