Media
September 16, 2026

How Investment Banks Should Measure AI Costs

By Samir Dutta, Co-founder and CEO of Farsight  ·  Published September 16, 2026

On August 21, OpenAI cut GPT-5.6 Sol API pricing by 20% for input and 33% for output for a promotional period. Three weeks earlier, it had cut Luna by 80% and Terra by 20%.

That caught my attention because investment banks can spend months evaluating an AI vendor while model pricing changes several times during the same procurement cycle. It also makes the true cost of AI in investment banking harder to compare using model rates alone.

When banks compare vendors, the quoted model rate can make one product look much cheaper than another. That number covers one part of what the bank pays. Retries, data retrieval from source files, automated tie-outs and validation checks, analyst or associate review, VP or Director review, and rework all add to the cost before anything reaches a client.

I think the better comparison starts with the finished work: the cost of producing a client-ready deliverable at the quality the firm expects.

Why token pricing understates the cost of AI

Published model prices vary widely. Here are 5 current examples:

Published API rates as of September 11, 2026. Billing conventions differ by provider, and prices can change.

The spread reflects both provider pricing and model tier. Banks should compare models on the same work and at the same quality standard.

The table shows the price of inference: what the provider charges for model input and output. Review time and the cost of finished work require separate measurement.

The billable token count for the same prompt can vary across models and providers. Reasoning can increase billable output, and long inputs can move requests into higher-priced tiers. Prompt design, caching, retrieval, provider selection, and the architecture around the model can reduce model spend.

Where human review adds to AI cost

An analyst or associate checks the work, fixes errors, and brings it up to the firm's standards. A VP may send it back with comments, and senior bankers review the analysis, sourcing, and language before anything reaches a client.

Weak sourcing means more checking. Analysis that does not tie to the underlying data has to be rebuilt, while a claim that does not fit the firm's view triggers another round of comments.

The cost also depends on the work:

Say a cheaper model saves $20 in generation cost and adds 30 minutes of associate review and 15 minutes of VP review. Now compare the $20 saving with the cost of those 45 minutes.

How to measure AI cost through approval

Fully loaded cost per approved deliverable gives banks a practical way to compare AI workflows.

The approval standard should match the work. For a pitch deck, it may include factual support, house style, and senior approval. Diligence work may require traceable sources. A financial model may need correct formulas, cross-footing, and references.

If you’re running a pilot, the most useful measures are:

  • Fully loaded cost per approved deliverable
  • Time to approval
  • Number of review cycles to approval
  • First-pass acceptance rate

Use the same measures before and during the pilot.

Why workflow design matters as model prices fall

Model prices will keep moving. Some savings are straightforward: reuse stable inputs, retrieve only the evidence needed for the task, and batch non-urgent work. Formatting, cross-footing, reference pulls, reconciliation, and rule-based checks usually have a clear expected result and should run consistently.

Other decisions depend on experience. A banker may need to decide whether a comparable belongs in the analysis, whether a source supports the claim, or whether the language fits the client.

3 questions to ask when evaluating an AI vendor

When evaluating an AI vendor or an internal pilot, I want answers to these questions:

  1. What does a finished deliverable cost? Include model use, retries, data retrieval, automated tie-outs and validation checks, and human review. Understand how provider pricing affects the total cost.
  2. Where does banker judgment enter the workflow? Identify the steps that require interpretation and the steps that can follow a defined process.
  3. If inference becomes 10 times cheaper, what happens to the economics? Recalculate the business case with model costs at a tenth of today's level.

Bank procurement moves slowly enough that a business case built today should still hold after another round of model price cuts.

How senior review becomes institutional knowledge

Senior redlines contain decisions the finished file no longer shows: claims the reviewer changed, comparables they kept, sources they trusted, and edits they made before approval.

Those decisions come from real engagements and years of review. Preserving them gives the next team a clearer set of standards to work from. As the same decisions recur, they can become firm preferences and workflow rules.

Farsight calls this the System of Judgment. It uses approved redlines and review decisions as context, so teams can apply the firm’s standards to new work.

Over time, that gives each new team a better starting point for the next engagement.