CapacityOS

Prove a VPP holds before someone deploys it.

Drop a new AI data centre, housing development, or EV depot onto the live Ontario Southwest Zone Sandbox, and watch whether a specific flexibility program can absorb it — real demand, real physics, a real optimizer.

18
owner agents
4
seasons tested
3
load types
waterloo-electric.vercel.app/sandbox
A live run in Ontario Southwest Zone Sandbox: a 20 MW data centre absorbed, result holds
Live product, real run
The national problem

Canada wants to build more. Growth increasingly arrives as electrical load.

Housing
New builds add heat pumps, EVs, and induction cooking to the grid, not just square footage.
EV charging
Depots and home chargers concentrate large, coincident loads at predictable hours.
Industry
Electrification of process heat and equipment turns fuel demand into grid demand.
Compute
AI data centres arrive as single loads the size of a small town — and run near-continuously.
…feeding one rising demand curve into a fixed local capacity limit.

A traditional interconnection study is slow, static, and answers only one scenario at a time. Meanwhile, flexibility programs are already real money: Ontario's Peak Perks VPP has over 100,000 enrolled homes today.

100,000+
homes enrolled in Ontario's Peak Perks VPP
Why not just batteries

Installed DER capacity is not deliverable VPP capacity.

Five things multiply — not add — to get from nameplate hardware to what a program can actually call on, in the hour that matters.

Physical capability×Participation×Incentive×Customer constraints×Timing=usable flexibility
Hover or tap a term above for a concrete example.
Where CapacityOS fits

We are not a capacity map, and we are not a DERMS.

Existing tools
Capacity visibility
Where is capacity?
Utility capacity maps and hosting-capacity tools tell you what headroom exists today.
CapacityOS
VPP design
What program design holds?
CapacityOS stress-tests a specific flexibility program against a specific new load, before anyone commits.
Existing tools
VPP operation
How do we run enrolled resources?
DERMS and VPP operating platforms enroll, dispatch, and settle real participants day to day.

CapacityOS sits in the design and stress-testing gap between knowing where capacity exists and operating enrolled resources day to day. It doesn't replace either side.

The product

Every one of these is a real, working part of the sandbox.

Ontario Southwest Zone Sandbox with a 7.3 MW overload, breaking, numbered hotspots over each real feature
1
Add demand
2
Edit VPP rules
3
Owner responses
4
Physical validation
5
Result
6
Fix It

Hover a number. Every screenshot on this site is from the live product, not a mockup.

The differentiator

When it breaks, CapacityOS doesn't just tell you. It tests what would fix it.

Every option shown is a verified rerun of the real day — never an LLM guess, never an estimate.

BEFORE — BREAKS
Before: breaks, 0.0 of 7.3 MW absorbed
APPLY → RERUN → VERIFIED
Fix It: two verified pathways, incentive change or capacity increase
Trust architecture

AI proposes. Physics constrains. Optimization dispatches.

LLM
Owner agents
Model willingness and price. Offer, decline, or revise.
Deterministic
Physical validator
Enforce battery SOC, EV deadlines, building comfort — every offer, checked.
Deterministic
OR-Tools optimizer
Decide the actual dispatch from what survived validation.
Derived
Result
Holds, partly holds, or breaks — traceable to every step above.

Bulk searches — Fix It, the season comparison — use the deterministic owner policy instead of live LLM calls. Repeated real calls would be slower and non-reproducible; a search needs the same conditions every time.

What's real, what's modeled

Ontario Southwest Zone Sandbox is not a digital twin. That's a feature of the brand, not an apology.

ObservedPublic IESO hourly reporting for the Southwest zone
DerivedThe hourly demand shape built from that reporting
Modeled234 device clusters, 18 owner agents, the 90 MW zone capacity
HypotheticalAny load you drag in — a scenario you're testing, not a real project
CalculatedEvery offer, dispatch, and result — always derived, never hand-tuned
What happens after the hackathon

The simulator doesn't change. Its inputs do.

Today
Sandbox
Public IESO demand data, a seeded synthetic DER population, an assumed zone capacity.
Next
Calibrated utility pilot
Swap in a utility's local constraints, real DER inventory, and the program rules actually being considered.
Then
Operational deployment
Hand the selected, stress-tested design to a real DERMS or utility program to enroll and run.
Who uses it, and the decision it answers

Not a persona. A question they need answered.

A utility asks:
“What participation target do we need to defer this upgrade?”
An aggregator asks:
“What incentive and resource mix should we deploy for this request?”
A consultant asks:
“Which program design survives the hardest historical scenario?”
A municipality asks:
“How much can flexibility contribute before infrastructure is still required?”
Before you launch it, three limits
  • ●One-zone capacity screening, not feeder- or device-level power flow.
  • ●Owner agents are modeled economic behavior, not calibrated predictions of real customers.
  • ●CapacityOS does not operate real resources — it's a design and stress-testing tool, not a DERMS.
Break a VPP before somebody deploys it.
Launch Ontario Southwest Zone Sandbox →
Built for AF Hacks: Growing Canada. A planning simulation, not a forecast or a utility-grade engineering study.
● hypothetical · ● modeled · ● derived · ● observed