↗Darshankumar Joshi
← All field notesFIELD NOTES / BLOG

DAF: what a fast agent needs beyond a fast model

AT A GLANCE

A personal engineering field note on Darshan’s agent framework, deterministic task graphs, and two small experiments in coordination latency—with the limits made explicit.

Two independent task cards converge on a finished notebook, drawn in ink on warm paper
A conceptual illustration of explicit dependencies, not benchmark data.

I built DAF to explore a practical question: when an AI workflow already has a stable plan, can explicit dependencies reduce the time spent coordinating it? A fast agent needs a clear execution contract as much as a capable model.

Explore DAF on GitHub — inspect the framework, study the architecture, and contribute a reproducible evaluation.

DAF is Darshj’s Agent Framework, also described as Darshan’s agent framework or Darshan’s agent. I publish as Darshankumar Joshi; readers searching for Darshan Kumar Joshi can find my work here.

Keep the reasoning; make the dependencies explicit

Consider a workflow with three tasks. One worker writes a slug-normalization function. Another writes a duration parser. A third reviews both outputs after they are ready. The first two tasks can run independently. The review must wait.

A directed acyclic graph, or DAG, expresses this without a conversation: two independent nodes feed a review node. The model still does the difficult coding and review. A deterministic scheduler manages readiness, limits concurrency, and records completion.

In the experimental DAF integration, GraphExecutor scheduled the work. DDAL, DAF Direct Agent Link, carried requests and results over a local TCP exchange. A bridge invoked actual signed-in CLI model workers. Sled persisted exact results, which were reopened and retrieved for dependent tasks.

This was an experimental adapter around DAF, not a built-in Codex execution mode. The public DAF repository and the local model integration are separate artifacts. Cloning public DAF does not, by itself, reproduce the complete measured setup.

Experiment one: eight paired ledger workflows

The ledger study used four held-out synthetic fixtures, repeated twice. Each workflow required two leaf workers followed by synthesis. The task included latest-version selection, approval filtering, signed ledger totals, weighted synthesis, and exact preservation of source nonces. Both arms used Astra at medium reasoning with a maximum of two leaf workers.

All eight native workflows and all eight DAF workflows passed. The native median wall time was 65.450 seconds; DAF’s was 18.109 seconds. The median of the eight within-pair native/DAF ratios was 3.65×, corresponding to a median paired latency reduction of 72.6%.

The clock included coordinator or process startup, inference, and result handling. Builds were excluded. Run order alternated between native-first and DAF-first, and all scored attempts were retained.

One revealing trace result: the two leaf intervals overlapped in all eight DAF runs, and in none of the eight native runs. The native coordinator scheduled those short workers sequentially despite having permission for two concurrent leaves. That was an observed behavior, not grounds for discarding its trials.

A later sensitivity check let the native coordinator perform synthesis itself after the two leaf workers. Those four runs all passed and had a 47.16-second median. They were later, unmatched runs, so they are not additional paired evidence. They do show why baseline topology matters: avoiding another delegation can reduce latency without adopting DAF.

Experiment two: actual code, with a corrected harness

The coding comparison recovered the original slugify and parse_duration requirements. Two workers generated the independent modules; a third reviewed them. Both arms used the same explicit task requirements, output ownership, read-only sandbox, no worker tools, and the same 187 independent behavioral cases for each coding output.

The corrected batch contained two alternating pairs:

  • Pair one: native 104.017 seconds; DAF 40.014 seconds.
  • Pair two: DAF 38.480 seconds; native 91.347 seconds.

Native passed both workflows, and DAF passed both. Every successful coding output passed the 187 behavioral cases. The median wall times were 97.682 and 39.247 seconds, respectively; the median paired ratio was 2.487×.

Both approaches showed overlapping leaf intervals here. The latency difference therefore cannot simply be explained as native workers never running in parallel.

The correction matters. An initial six-run batch had unequal output contracts and was excluded in full. It was not counted as a native coding failure. The original review context also lacked dependency task specifications; the corrected comparison supplied those specifications equally to both arms. A historical 37.680-second development run was kept separate and was not used in the ratios.

Timing included startup, all model calls, transport, persistence, and result materialization. Setup, build, trace extraction, and behavioral testing were excluded equally. The review’s execution was verified, but its prose did not have a deterministic semantic correctness oracle.

What the results cannot tell us

Four unique synthetic fixtures and two coding pairs are small samples. They do not establish broad statistical significance, stronger reasoning, or superiority across software development. Service load, cache behavior, network conditions, trial order, and context construction can affect wall time.

Native subagents and standalone CLI workers did not receive identical internal system contexts. Encrypted native child messages also limited independent byte-for-byte verification of prompt forwarding. Matching the exposed model, requirements, and checks did not eliminate every architectural difference.

The experiments do not isolate a speed benefit from DDAL or Sled. A substantial part of the difference may come from replacing model-driven coordination and its context overhead with explicit scheduling. Stronger native baselines could narrow that difference.

I am not claiming billed-token or dollar savings. Usage scopes were retained, but native parent and child accounting could not safely be added because inclusion semantics were unclear. No invoice-level comparison was available.

The 3.65× ledger result and the 2.487× coding result belong to different studies. They should not be pooled into one headline statistic.

The hard part is the execution contract

A scheduler that launches workers is easy to demonstrate. A scheduler that handles interrupted work safely needs a much stronger contract.

Completed results should be tied to the task, input snapshot, model configuration, and permitted capabilities. A changed input should invalidate reuse. Dependencies should be released only after an acceptable result has been committed. Concurrency limits should control resource admission rather than merely label a plan as parallel.

Failure, cancellation, and an uncertain outcome are different states. A lost connection after an external action does not mean the action failed. Repeating a draft-generation call and repeating a payment are not equivalent operations. Recovery needs operation-specific retry and reconciliation rules.

The prototype demonstrated reuse of completed nodes. It did not establish exactly-once remote execution or arbitrary hard-crash recovery. Killing a local process also does not prove the remote computation stopped or stopped billing.

Permissions remain part of execution. A graph must not grant a worker capabilities denied by the user or platform. Tool-enabled coding needs isolated workspaces, clear output ownership, and normal conflict review. Deterministic coordination is useful only if it preserves these protections.

Where this approach fits

An explicit graph is a plausible fit when the plan is stable, outputs are well defined, independent tasks exist, and correctness can be checked. Independent module proposals, bounded analysis, and structured synthesis are useful candidates.

Open-ended investigation is different. The next task may depend on what a model discovers, and a human may change direction. A model-driven coordinator can be valuable there. The practical question is which portions of a workflow have become predictable enough to schedule explicitly.

The next benchmark should try to disprove the advantage

The next evaluation should be independent, larger, and fixed before the results are known. It should include mostly serial work, deeper dependencies, strong native coordinator baselines, failures, and safe recovery tests. Run order should be balanced; incomplete runs should remain visible; model settings, task inputs, output contracts, and validation should be pinned.

A proposed design is 30 distinct tasks with five paired repetitions each—300 workflow executions across two arms. That is a proposal, not work already completed. A smaller pilot should first establish feasibility and comparable instrumentation.

The question I want answered is precise: for which bounded workflows can explicit dependency scheduling reduce end-to-end latency without reducing correctness or weakening user control?

These two experiments provide a reason to test that question carefully. They do not settle it.

Source and availability

DAF source repository

The figures above come from retained local ledger and corrected coding reports. The benchmark adapter and detailed replication package are not represented here as publicly downloadable. This article describes experimental work and does not imply OpenAI endorsement or a native integration agreement.

Illustrations are conceptual; they do not encode benchmark data.

Frequently asked questions

What is DAF?

DAF is Darshj’s Agent Framework. The experiment used explicit task dependencies to coordinate model workers while preserving their coding and review work.

Does DAF make every AI agent 2.5 times faster?

No. The corrected coding study had only two paired workflows and a median paired wall-time ratio of 2.487×. It does not establish universal speed, stronger reasoning, or token and cost savings.

Can I reproduce the benchmark by cloning the public repository?

Not by cloning alone. The public DAF framework and the local experimental model adapter are separate artifacts; the complete measured integration is not represented as a public download.

Keep exploring field notes ↗
Let’s talk it through ↗
Let’s chat