Hinweis
Für den Zugriff auf diese Seite ist eine Autorisierung erforderlich. Sie können versuchen, sich anzumelden oder das Verzeichnis zu wechseln.
Für den Zugriff auf diese Seite ist eine Autorisierung erforderlich. Sie können versuchen, das Verzeichnis zu wechseln.
Note
Hilfen zum Self-Hosting für OpenAI Responses-Endpunkte in .NET kommen bald.
Note
Hilfsprogramme für Self-Hosting von OpenAI-Responses-Endpunkten sind derzeit für Go nicht verfügbar.
Verwenden Sie agent-framework-hosting-responses, um Anfragen und Antworten im Format von OpenAI Responses über einen Endpunkt zu konvertieren, der Ihrer Anwendung gehört. Ihr Server wählt das Webframework, die Route, Authentifizierung, Autorisierung, Anforderungsoptionen und Sitzungsspeicher aus.
pip install --pre agent-framework agent-framework-foundry agent-framework-hosting agent-framework-hosting-responses azure-identity
Das FastAPI-Beispiel ist eine Implementierung. Die gleichen Helfer arbeiten mit Django, Flask, Starlette, Azure Functions oder einem anderen Framework.
Hosten eines Agent-Endpunkts
Dieses Beispiel konvertiert die Anfrage in Ausführungswerte des Agent Framework, wendet eine anwendungsdefinierte Zulassungsliste für Optionen an und speichert die aktualisierte Sitzung unter der neu erstellten Antwort-ID.
app = FastAPI()
state = AgentState(create_agent)
ALLOWED_REQUEST_OPTIONS = frozenset({"max_tokens", "reasoning"})
@app.post("/responses", response_model=None)
async def responses(body: dict[str, Any] = Body(...)) -> JSONResponse | StreamingResponse: # noqa: B008
"""Handle one OpenAI Responses-shaped request."""
try:
run = responses_to_run(body)
except ValueError as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
session_id, is_conversation_id = responses_session_id(body)
conversation_id = session_id if is_conversation_id else None
response_id = create_response_id()
# App-specific policy: allow only the request options this route is willing
# to honor. This denies tools, tool_choice, deployment/persistence fields,
# and all other caller-supplied options by default. Your app decides which
# options are allowed, altered, or denied.
options = {key: value for key, value in run["options"].items() if key in ALLOWED_REQUEST_OPTIONS}
options["reasoning"] = {"effort": "medium", "summary": "auto"}
options_for_run = cast(Any, options)
target = await state.get_target()
lookup_id = session_id or response_id
# An unknown `conversation_id` becomes a new session here. Production apps
# can choose to require a separate "create conversation" API instead.
session = await state.get_or_create_session(lookup_id)
if run["stream"]:
stream = target.run(
run["messages"],
stream=True,
session=session,
options=options_for_run,
)
if not isinstance(stream, ResponseStream):
raise HTTPException(status_code=500, detail="agent did not return a response stream")
async def stream_events() -> AsyncIterator[str]:
async for event in responses_from_streaming_run(
stream,
response_id=response_id,
conversation_id=conversation_id,
):
yield event
# `agent.run(..., stream=True)` updates the session while the stream
# is consumed/finalized. Persist the selected continuation only
# after finalization.
if conversation_id is not None:
# A stable conversation id is a mutable head. Apps must ensure
# only one caller advances it at a time; AgentState does not
# serialize concurrent runs for the same id.
await state.set_session(conversation_id, session)
else:
await state.set_session(response_id, session)
return StreamingResponse(
stream_events(),
media_type="text/event-stream",
)
result = await target.run(
run["messages"],
session=session,
options=options_for_run,
)
# `agent.run(...)` updates the session. Persist the selected continuation
# only after the run completes.
if conversation_id is not None:
# Preserve sequential conversation continuity. Production apps must
# provide their own per-conversation single-writer coordination.
AgentState ermittelt das Ziel und lädt oder erstellt eine Sitzung. Speichern Sie die Sitzung nach der Ausführung oder nach Abschluss einer Streamingausführung, da die Ausführung sie aktualisiert.
Die vollständige Anwendung, einschließlich der Agentdefinition und der Positivliste für Anforderungsoptionen, finden Sie im Beispiel „Local Responses“.
Hosten eines Workflowendpunkts
WorkflowState löst den Workflow auf, ihre Anwendung besitzt jedoch den Prüfpunktspeicher und die Zuordnung von einer Antwort-ID zu einem Prüfpunkt. Dieses Beispiel stellt den von einem autorisierten previous_response_id ausgewählten Checkpoint wieder her und speichert anschließend einen Cursor für die nächste Antwort.
app = FastAPI()
state = WorkflowState(workflow_builder, cache_target=False)
@app.post("/responses", response_model=None)
async def responses(body: dict[str, Any] = Body(...)) -> JSONResponse: # noqa: B008
"""Handle one OpenAI Responses-shaped request for the workflow."""
try:
run = responses_to_run(body)
except ValueError as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
# This sample demonstrates only Responses `previous_response_id`
# continuation, so reject `conversation_id` instead of treating it as a
# checkpoint cursor.
previous_response_id, is_conversation_id = responses_session_id(body)
if is_conversation_id:
raise HTTPException(
status_code=400,
detail="This server supports previous_response_id continuation only; conversation_id is not implemented.",
)
response_id = create_response_id()
target = await state.get_target()
if previous_response_id and (checkpoint_cursor := checkpoint_cursor_store.get(previous_response_id)) is not None:
# Restore first. Workflow.run does not allow `message` and
# `checkpoint_id` in the same call.
await target.run(
checkpoint_id=checkpoint_cursor["checkpoint_id"],
checkpoint_storage=checkpoint_storage_for(checkpoint_cursor["storage_id"]),
)
storage_id = response_id
checkpoint_storage = checkpoint_storage_for(storage_id)
result = await target.run(
message=workflow_prompt_from_messages(run["messages"]),
checkpoint_storage=checkpoint_storage,
)
latest = await checkpoint_storage.get_latest(workflow_name=target.name)
if latest is not None:
# Responses `previous_response_id` can point to any response id. Store
# the current response id as the cursor for this workflow continuation.
cursor = CheckpointCursor(checkpoint_id=latest.checkpoint_id, storage_id=storage_id)
checkpoint_cursor_store.set_many({response_id: cursor})
return JSONResponse(
responses_from_run(
response_from_workflow_result(result),
response_id=response_id,
)
)
Die dateibasierte Speicherung des Beispiels ist für die lokale Entwicklung vorgesehen. Verwenden Sie persistenten Speicher, wenn Replikate neu gestartet oder horizontal skaliert werden können.
Important
Behandeln previous_response_id und conversation_id als nicht vertrauenswürdige Eingabe. Authentifizieren und autorisieren Sie den Anrufer, bevor Sie entweder die ID zum Laden oder Speichern einer Sitzung oder eines Prüfpunkts verwenden.
Informationen zum breiteren Drahtformat finden Sie unter OpenAI-kompatible Endpunkte.
Nächste Schritte
Gehen Sie tiefer: