Skip to main content

AI Agent Testing — Status 2026-07-14

Delivery / Program Management — periodic status report.

Summary

initiative-timeline-mapper ran in squad mode (2026-07-14) across the four DRI-sized Chatbot initiatives. This initiative (P3 in the roadmap Priority table) is off-track against 2026-Q3: under the current priority order and staffing, it is projected to finish after the quarter ends (p50 ~2026-10-04, already past 2026-09-30). This is a sequencing/capacity conflict, not an estimating error — the 37.5 md total itself is DRI-confirmed and plausible.

Highlights

  • Effort is DRI-confirmed: 37.5 md (BE 16.5 · FE 13 · QA 8), as-of 2026-07-09, medium-high confidence.
  • RFC review verdict is PROCEED — an agent can execute the full 13-chunk plan once a task breakdown exists; the only design gap (a missing Figma frame) blocks pixel-faithful work on one chunk only.

Lowlights & blockers

  • No task breakdown exists yet. Sprint 1's own outcome target 4 is to generate one for this initiative — until it lands, the phase split and sequencing below the total are unconfirmed, so this forecast is Low confidence.
  • Farras (BE) cannot start this initiative's BE prep until he clears Autonomous AI Agent (P1) — not before ~2026-07-27.
  • Baghiz (FE) is the real constraint. He is queued: unsized Google Calendar work → AI Agent Impact Report (P2, 15.5 md) → this initiative (13 md) — third in line. Under the P2 forecast above, he cannot start this initiative's FE before ~2026-08-24, and 13 md at his available capacity does not finish before ~2026-09-20.
  • External, no committed date: the AI platform squad's shared batch shadow-inference capacity (TPM/RPM limits) blocks QA execution outright — the RFC review names this as the build gate. There is no committed date from that squad, so the QA-complete date (and therefore the whole forecast) is a floor, not a ceiling.

Risk & dependency changes

  • New — quarter-fit conflict (not previously recorded): under the roadmap's P1→P2→P3 priority order, this initiative's real FE/QA start is pushed to late August/September purely by queue position, before the external AI-platform capacity gate is even considered. Recorded here per the timeline-mapper's persistence doctrine (a notable priority/timeline conflict is institutional memory, not just a derived date) — see also ../../delivery/roadmap.md ## Program-level risks.

Metrics

  • Progress: 0 / 37.5 man-days (0%) — planning-stage only this sprint (generating the task breakdown).
  • Forecast: p50 2026-10-04 · p85 2026-10-25 or later (basis: judgment, Low confidence — no task breakdown, external gate with no ETA).

Next period

  • Generate the task breakdown (Sprint 1 outcome target 4) to raise this forecast's confidence.
  • Leadership decision needed: accept the Q4 slip, add a second FE engineer to unblock Baghiz's queue, or reprioritize this initiative ahead of Google Calendar / AI Agent Impact Report.
  • Push the AI platform squad for a committed batch-capacity date — the single largest unknown.