My view is that AI should be introduced into transfer pricing (TP) processes in small, independently verifiable steps.

But the real challenge is deciding where one step ends and the next begins.

Why does this matter?

The LogiConBench study (ICLR 2026), involving researchers from the University of Amsterdam, provides interesting insights into the logical consistency of AI models.

When the number of interconnected logical statements increased from two to five, GPT-5 accuracy in zero-shot testing changed:

  • From 90% to 87.5% when checking whether one proposed combination was logically consistent.
  • From 38.3% to just 7.3% when identifying the complete set of logically consistent combinations.

These are two different tasks. The second figure measures fully correct answers: missing even one valid combination counts as an incorrect result.

A simple transfer pricing example

Consider a simplified TP calculation involving five pricing decisions: choosing between two base indices, adjusting for logistics costs, accounting for handling costs separately, adjusting for financing costs, and selecting between two margin treatments.

Five binary decisions create 32 theoretical combinations.

The challenge is not generating 32 combinations, but identifying which ones are valid under all relevant conditions.

For example, if handling costs are already included in logistics, adding them separately would result in double-counting.

The researchers found that AI models often missed valid combinations, repeated others, or failed to provide a complete answer.

Three further observations from the study:

  • Fine-tuning: Training smaller models improved results somewhat, but did not solve the problem.
  • Wording and prompting: Rephrasing statements made little difference. Better prompts helped, but did not solve the problem.
  • Logical graphs: Models received statements derived from graphs, not the full graphs. Whether access to the full graphs could improve performance remains an open question.

Of course, the study tested logical reasoning, not TP calculations. But it highlights an important difference between checking one proposed solution and identifying all valid alternatives.

The practical challenge for TP automation

I see two major difficulties in applying a step-by-step approach.

The first is identifying where to draw the boundaries between steps.

A task may involve a simple set of alternatives, or it may require several interconnected logical decisions, where one choice affects the options available at the next stage.

The distinction is not always obvious.

To identify these boundaries, TP specialists need a deep understanding of the TP process, its underlying logic, and how to translate that logic into clearly defined tasks for AI.

The second difficulty is connecting the individual steps.

A TP process is not always a linear sequence of decisions.

At certain stages, a TP specialist must decide which alternatives should be explored further. New findings may require revisiting earlier assumptions, sometimes going back several steps — and doing so more than once.

This makes the design of the overall workflow just as important as the reliability of each individual AI task.

Where does this leave us?

The challenge is not only to verify whether an AI-generated answer is correct, but also to ensure that no relevant alternatives have been missed.

This may require combining AI with structured rules, systematic search, and independent verification.

The key challenge for the future of TP automation is deciding which parts of the process can stand alone, how they should be connected, and where independent verification and professional judgment are essential.


Source: LogiConBench: Benchmarking Logical Consistencies of LLMs (ICLR 2026)⁠

Leave a Reply