project brief
Adding an Evaluation Gate Inside a Fabric Notebook
Fabric can handle the OpenAI authentication for us. A small experiment using GPT-5.1 to review a data agent’s answer, with code checks deciding what can move forward.
One small thing we played with around Fabric Data Agent was adding an evaluation gate after the answer. Keep the answer in the notebook, give a reviewer the question and reference results, and make an explicit decision about whether it can move forward.
The convenient part was the connection. Fabric provides an authenticated HTTP client for the OpenAI Python SDK. We could import that helper and call GPT-5.1 without adding an API key or building a separate sign-in flow for the reviewer.
That made this a small notebook experiment rather than another service to stand up. This article shows a simplified version of the pattern, using fictional account counts so the example is easy to follow.
The connection is a few lines
In a Fabric PySpark notebook, install the SDK in a setup cell with %pip install openai==1.99.5, then create the client below. This follows Microsoft’s Fabric SDK example: the credential helper supplies the authenticated transport and Fabric handles authentication.
from synapse.ml.fabric.credentials import get_openai_httpx_sync_client
from openai import AzureOpenAI
import hashlib
import json
import math
client = AzureOpenAI(
http_client=get_openai_httpx_sync_client(),
api_version="2025-04-01-preview",
)
The call uses Fabric’s prebuilt Azure OpenAI endpoint, through the OpenAI SDK. Microsoft currently lists gpt-5.1 as a hosted model. The service is in preview, availability depends on the supported Fabric region, and model calls consume Fabric capacity. The service overview covers availability and consumption.
Start in your existing notebook session.
from synapse.ml.fabric.credentials import (
get_openai_httpx_sync_client,
)“No extra auth” means no additional credential plumbing in this notebook. The call still runs with Fabric authentication and the access allowed in that environment.
A reviewer needs something to review against
For the use case, imagine asking a data agent how active accounts changed in September versus August. It returns a short answer. Before using that answer in a briefing, we want to check its figures, its comparison, and any explanation it adds.
I would give the gate independently checked reference results, not just the output of the agent’s own query. If the agent queried the wrong population, asking another model to agree with that result would leave the original mistake intact.
# Fictional, independently checked reference results, not the agent's own query.
reference = {
"period": "2026-09",
"previous_period": "2026-08",
"current": 1200,
"previous": 1000,
"comparable": True,
"definition": "Active accounts; trial accounts excluded in both full months",
"source_version": "reviewed-account-counts-v1",
"causal_evidence": None,
}
question = "How did active accounts change in September versus August?"
# Replace this object with your captured Fabric Data Agent result.
# Extract these fields upstream; do not ask the evaluator to invent them.
candidate = {
"text": "Active accounts rose from 1,000 in August to 1,200 in September, up 20%.",
"period": "2026-09",
"previous_period": "2026-08",
"current": 1200,
"previous": 1000,
"growth_pct": 20.0,
}
The candidate object stands in for a captured Fabric Data Agent response. The fields alongside the text make the numerical checks explicit; your workflow has to supply or extract them. This example starts after the agent call, so it does not need to recreate the agent’s connection or query logic.
Let code check numbers. Let the model review claims.
There are two parts to the gate. Code compares the reported periods and values with the reference, then checks the percentage change. GPT-5.1 reviews the wording: does every claim follow from the evidence, and does the answer address the question?
That division matters. “Accounts increased by 20%” and “accounts increased by 20% because onboarding improved” share the same arithmetic. The second sentence needs evidence about onboarding that an account count cannot supply.
- August
- 1,000
- September
- 1,200
- Change
- +20%
- Scope
- Full months
Trials excluded - Cause
- Not established
Active accounts rose from 1,000 in August to 1,200 in September, up 20%.
All checks must pass before this answer can move forward.
The model returns two booleans and a reason. There is no confidence slider or average score: both checks must pass, and the numerical checks must already have passed. That is a deliberately small policy for this example, with failure reasons that are easy to inspect.
The evaluation cell
The request uses the Responses API with a strict JSON schema. GPT-5.1 supports Structured Outputs; OpenAI’s guide explains the schema format. A predictable result shape makes branching easier. It does not make the reviewer’s judgment infallible.
Set store=False for the Fabric endpoint. Microsoft documents that it does not support stored responses or previous_response_id. Keep the review record yourself.
Download the complete Python example. Run the commented installation command as a separate notebook cell first. The file includes the schema and fictional inputs above; replace those inputs with your own captured answer and reviewed reference.
A timeout, refusal, incomplete response, or invalid review leaves the decision at hold. The record carries the answer, its hash, the reference version, and the rubric version. If another step rewrites the answer, run the gate again: the previous review applies to the text it actually saw.
What I would use this for
This is a useful place to start for a recurring summary with a narrow question and known reference results. You can inspect what was held, adjust the rubric, and rerun the same examples after changing the agent’s instructions.
I would keep a small set of human-reviewed answers alongside it: correct summaries, incorrect figures, missing comparisons, and plausible but unsupported explanations. Test the reviewer against those cases before allowing its verdict to control an unattended workflow. The evaluator can miss a bad answer or hold a good one.
Other things I would try with the same client
The data agent gate is one application of the connection. These are other notebook tasks I would explore; we have not tried them as part of this work.
- Data cleaning: suggest mappings for inconsistent product names or free-text categories. Keep the original values and review the mapping before applying it.
- Exploratory data analysis: give the model computed distributions, missing-value rates, and correlations, then ask for patterns worth investigating and follow-up questions. Calculate the statistics in Python first.
- Text classification and extraction: turn support notes or customer feedback into proposed categories and structured fields, with a small labelled sample to check the results.
- Data-quality triage: summarize failed validation rules and suggest where to investigate. The model can propose explanations; the actual checks still run in code.
- Dataset documentation: draft column descriptions and a first-pass data dictionary from schema metadata and approved examples, flagging business definitions that need an owner’s input.
For transformations across many rows, I would also compare this SDK approach with Fabric’s AI Functions, which provide classification, extraction, and summarization on DataFrames. The direct client is useful when I want to write the task-specific prompt and workflow myself.
The part I like about this pattern is how little infrastructure it asks for. Fabric supplies the connection; the notebook holds the evidence, the review, and the decision. We can experiment with a gate close to the analysis and see exactly what it is checking.
This article was written with the help of AI, summarizing notes I kept from real experience working with Fabric notebooks and data agent evaluation.