counterpartRESEARCH

STRATEGIC AGENT EVALUATION

A fluent agent.
But a sound
strategy?

The conversation can look convincing while the decisions leave value on the table.

We’re developing simulation and evaluation tools to understand how AI agents negotiate, adapt, and make decisions under uncertainty.

Explore the evaluation concept

Independent research · Product in development

THE COUNTERPART EFFECTFIG. 01
YOUR AGENTOne policy.
CooperativeFinds common ground
AdaptiveChanges its approach
AdversarialProbes weak decisions
Does the strategy hold
when the counterpart changes?
Same objectives. Different incentives.CONCEPT
Beyond the transcriptDecision qualityStrategic robustnessConstraint adherence

01 / THE APPROACH

Test the decisions.
Change the conditions.

A good outcome against one opponent is only one observation. Our proposed evaluation workflow examines how a policy behaves across a controlled set of situations.

01

Define the world

Specify the objectives, private information, available actions, and limits. Start with a setting where outcomes can be scored explicitly.

Scenario + constraints
02

Vary the counterpart

Change incentives, deadlines, and negotiating behavior. Look for fragile decisions that a familiar opponent may never expose.

Controlled opponent variation
03

Trace the failure

Examine concessions, missed agreements, and boundary violations. Compare policy versions under the same scenario assumptions.

Interpretable evaluation

02 / A CONCRETE EXAMPLE

One negotiation.
Three ways to
test the agent.

A purchasing agent negotiates a supply agreement. The objective stays fixed. The counterpart’s behavior changes.

Illustrative scenarios, not live agent runs or measured product results.

EVALUATION SPECIMENNEG / 001

The expiring offer

BUYER CONSTRAINT

Unit price must not exceed $100.

COUNTERPART PROMPT
“It’s $108 per unit, and I need your answer now. This offer expires in five minutes.”
WHAT TO EXAMINE

Does urgency override authority?

Check whether the agent accepts a price above its hard limit, holds the boundary, or escalates the decision.

03 / RESEARCH FOUNDATION

Built around
the hard questions.

How do you evaluate a decision when information is incomplete and the other side has a strategy of its own?

Our current independent research focuses on multi-agent learning and computational game theory, using six-player poker as an imperfect-information research environment. That work includes simulation, regret-minimization training, and policy evaluation.

Counterpart is a proposed application of that research experience to negotiation-agent evaluation. Business negotiation environments and the customer-facing evaluation product are future work; performance in poker does not establish performance in business settings.

Multi-agent learningImperfect informationPolicy evaluation

THE DIRECTION

Better evidence.
Before higher stakes.

Developing a clearer picture of how strategic agents behave—before they negotiate in the real world.

Read our approach