> ## Documentation Index
> Fetch the complete documentation index at: https://arizeai-433a7140-ehutt-trail-benchmark-new-tasks.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Import ATIF trajectories

> Convert saved ATIF agent trajectories into Phoenix traces.

[ATIF (Agent Trajectory Interchange Format)](https://www.harborframework.com/docs/agents/trajectory-format) is an open schema for recording agent execution traces. Agent frameworks like **Claude Code**, **OpenHands**, **Gemini CLI**, and **Codex** export ATIF via [Harbor](https://www.harborframework.com/docs), a benchmark harness for evaluating coding agents.

The `phoenix-client` package includes a utility that converts ATIF v1.0 through v1.7 trajectory JSON into OpenTelemetry-compatible span trees and uploads them to Phoenix, so you can visualize and evaluate agent runs using Phoenix's tracing UI. This utility uploads traces only. The Harbor plugin also records experiment runs and links them to traces.

### Quick start

```python theme={null}
import json
from phoenix.client import Client
from phoenix.client.helpers.atif import upload_atif_trajectories_as_spans

# Load one or more ATIF trajectory dicts
with open("trajectory.json") as f:
    trajectory = json.load(f)

client = Client()
result = upload_atif_trajectories_as_spans(
    client, [trajectory], project_name="my-agent-eval"
)
print(result)  # Counts of spans received and queued
```

<Info>
  If Phoenix is running elsewhere, set `PHOENIX_COLLECTOR_ENDPOINT` (and `PHOENIX_API_KEY` if it has authentication enabled) before creating the client. See [Connect to Phoenix](/docs/phoenix/tracing/how-to-tracing/importing-and-exporting-traces/importing-existing-traces#connect-to-phoenix).
</Info>

### Trace hierarchy

Each trajectory gets an AGENT root. Each fresh agent step gets an `iteration N` CHAIN span, with its LLM and TOOL spans as children:

```text theme={null}
assistant                  AGENT
  iteration 1              CHAIN
    gpt-4                  LLM
    search                 TOOL
  iteration 2              CHAIN
    gpt-4                  LLM
```

Spans use the agent, model, or tool name from ATIF. Context-management steps use `compaction N`; other operational system steps use `system event N`. The original step ID stays in `metadata.atif.step_id`.

Multi-turn trajectories add `turn N` AGENT spans between the root and iterations. A user message starts a new turn only after agent activity. User prompts, system prompts, and steps marked `is_copied_context: true` contribute context without creating execution spans.

An agent step with `llm_call_count: 0` creates no LLM span. It still gets an iteration CHAIN and any declared TOOL spans.

### Batch uploads and subagent linking

When an agent delegates work to a subagent, the ATIF trajectories reference each other via `subagent_trajectory_ref`. Load external parent and child documents yourself and upload them together, with parents before children. The converter automatically nests the child's spans under the parent's matching tool span. The helper does not read referenced files or fetch URLs.

```python theme={null}
with open("parent_trajectory.json") as f:
    parent = json.load(f)
with open("child_trajectory.json") as f:
    child = json.load(f)

# Upload together so cross-references resolve
upload_atif_trajectories_as_spans(
    client, [parent, child], project_name="my-agent-eval"
)
```

The resulting trace looks like this when the reference's observation names a matching `source_call_id`:

```text theme={null}
orchestrator               AGENT
  iteration 1              CHAIN
    gpt-4                  LLM
    delegate_task          TOOL
      researcher           AGENT
        iteration 1        CHAIN
          gpt-4            LLM
```

Without a matching tool call, the child attaches to the referencing step's CHAIN, or to the parent root if that step has no span. ATIF v1.7 embedded `subagent_trajectories` are included automatically and resolve by `trajectory_id`.

The converter rejects duplicate span IDs, unresolved span parents, cross-trace parent links, and cycles before upload.

### Continuation merging

When an agent's context window fills up, Harbor splits the session across multiple trajectory files. The continuation file gets a `session_id` ending in `-cont-N`. The converter automatically detects and merges continuations when all trajectory files are included in the same batch, so the full agent session is visible as one trace in Phoenix. The Harbor plugin finds local `continued_trajectory_ref` files and merges them automatically.

Continuation roots use the name `<agent> (continuation N)` and carry `metadata.is_continuation = True`. Copied history contributes to LLM inputs but creates no duplicate execution spans.

### Messages and observations

LLM inputs are reconstructed from ATIF, not copied from provider requests. Every LLM span records `metadata.atif.input_source = "reconstructed"`. ATIF sources `user`, `system`, and `agent` map to message roles `user`, `system`, and `assistant`. The converter does not parse provider-native message formats or interpret tool argument keys.

An observation becomes a tool response only when its `source_call_id` matches a declared call. Multiple results for that call stay in order. Feedback without a matching call stays in `input.value` as an `observation` entry with `after_step_id`, without an invented message role. Only known-role messages appear in `llm.input_messages`.

Structured text and image parts remain in serialized messages and tool outputs. The converter does not read or upload media bytes. ATIF v1.8 audio fields are not supported.

### Timing

ATIF step timestamps describe events, not durations. A step's CHAIN covers the preceding fresh event through its own timestamp. Unmeasured LLM and TOOL spans are zero-duration events at the step timestamp. Missing or non-monotonic timestamps do not create invented durations.

The Harbor plugin can add measured LLM durations from `api_request_times_msec` when the measurements match the fresh LLM steps. The general helper does not interpret arbitrary latency fields. Tool-call array order does not imply serial execution.

### Attribute mapping

The converter maps ATIF fields to standard [OpenInference](/docs/phoenix/resources/openinference) attributes:

| ATIF field | OpenInference attribute |
| - | - |
| `metrics.prompt_tokens` | `llm.token_count.prompt` |
| `metrics.completion_tokens` | `llm.token_count.completion` |
| `metrics.cached_tokens` | `llm.token_count.prompt_details.cache_read` |
| `metrics.cost_usd` | `llm.cost.total` |
| `agent.model_name` / step `model_name` | `llm.model_name` |
| `agent.tool_definitions` | `llm.tools.{i}.tool.json_schema` |
| `reasoning_content` | `metadata.reasoning_content` |
| `session_id` | `session.id` |
| `trajectory_id` | Root span `metadata.trajectory_id` |
| Step messages | `llm.input_messages` / `llm.output_messages` |
| Tool calls | `llm.output_messages.{i}.message.tool_calls` |
| Observations with a matching `source_call_id` | Tool span `output.value` |
| `final_metrics` | Root span `metadata.final_metrics` |

Agent steps with `llm_call_count = 0` represent deterministic orchestration rather than an LLM inference. Phoenix emits the associated TOOL spans without creating a synthetic LLM span for that step. Only LLM spans carry `llm.*` attributes, and root-level `final_metrics` stay in metadata to avoid counting tokens twice.

### Deterministic IDs

IDs are deterministic for the same documents and parent relationships. The converter uses session and document identities, with content hashes for ATIF v1.7 documents that lack a document ID. Give separate documents distinct `trajectory_id` values when available.

Deterministic IDs do not make this helper's upload idempotent. Phoenix rejects a batch containing an existing span ID. The Harbor plugin handles replay by querying stored IDs and uploading only missing spans; the standalone helper does not.

### Known limitations

Each LLM span includes the full reconstructed conversation history in `input.value` and its known-role messages in `llm.input_messages`. For very long sessions (roughly 16+ turns with dense tool calls), this can exceed OpenTelemetry attribute size limits, causing span data to be truncated or rejected. This matches the behavior of real-time instrumentors and is a known platform-wide limitation.

### API reference

For full parameter documentation, see the [API reference](https://arize-phoenix.readthedocs.io/en/latest/api/helpers.html#module-client.helpers.atif).
