Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
147456c2a2 |
@@ -60,30 +60,13 @@ that subtraction into the add-on, where it is done once and tested, and replaces
|
||||
|---|---|---|
|
||||
| `meter_source` | `off` | `off` keeps `meter_entity`. `ha_dsmr` subscribes to the DSMR integration over the HA WebSocket; `mqtt_p1` reads a topic |
|
||||
| `meter_phases` | 1 | 1 or 3. Must match the telegram, or every telegram is rejected and logged |
|
||||
| `meter_max_age_s` | 30 | Beyond this the reading is stale and grid power reads as *missing*. On its own it does **not** command 0 W — see the timing note below. It is also the longest a reading is held forward into the 15-minute average |
|
||||
| `meter_max_age_s` | 30 | Beyond this the reading is stale: grid power reads as *missing*, and the existing failsafe commands 0 W |
|
||||
| `meter_mqtt_topic` | | `mqtt_p1` only |
|
||||
| `p1_import_entity` | | The **unsigned** consumption sensor. Do not point this at a signed template |
|
||||
| `p1_export_entity` | | The **unsigned** injection sensor |
|
||||
| `p1_phase_import_entities` | `[]` | L1..L3, in order. Needed for the capacity-tariff peak on a three-phase connection |
|
||||
| `p1_phase_export_entities` | `[]` | L1..L3, in order |
|
||||
|
||||
#### How long a dead meter takes to reach 0 W
|
||||
|
||||
`meter_max_age_s` and `stale_input_s` **stack**. They are two different clocks
|
||||
and neither one is the whole answer:
|
||||
|
||||
| step | option | default |
|
||||
|---|---|---|
|
||||
| telegrams stop, P1 sample goes stale, grid power starts reading *missing* | `meter_max_age_s` | 30 s |
|
||||
| inputs have been missing long enough for the loop to command 0 W | `stale_input_s` | 15 s |
|
||||
| **total, meter death → 0 W commanded by this add-on** | | **45 s** |
|
||||
|
||||
So in P1 mode `stale_input_s` is *not* "how long inputs may be missing before
|
||||
commanding 0 W" measured from the meter dying — it is measured from the moment
|
||||
the P1 sample already went stale. Size the pair together: the ESP32's own
|
||||
watchdog commands 0 W after ~30 s of silence from this add-on regardless, and
|
||||
that layer is unaffected by either option.
|
||||
|
||||
There is **no fallback to an inverter-side power figure**, deliberately. The
|
||||
inverter's own AC power tracks its battery almost perfectly and the real meter
|
||||
hardly at all, so a controller that failed over to it would be regulating
|
||||
@@ -119,18 +102,6 @@ emits nothing, which is indistinguishable — to anything watching the value —
|
||||
a meter that has died. Watching the age instead separates the two: it climbs when
|
||||
telegrams stop and resets when they arrive, whatever the reading says.
|
||||
|
||||
The entity is only created when `meter_source` is not `off`. With P1 ingestion
|
||||
disabled there is nothing feeding it, and an age sensor climbing with no ingester
|
||||
behind it would trip the firmware watchdog on a system that is working fine.
|
||||
|
||||
> **Known limit, `mqtt_p1` only.** The age measures *arrival*, not change. On the
|
||||
> `ha_dsmr` path that is exactly right: a frozen meter emits no `state_changed`,
|
||||
> so nothing arrives and the age climbs. On the MQTT path a bridge that is stuck
|
||||
> republishing its last telegram keeps arriving, so the age stays near zero and a
|
||||
> frozen meter still looks fresh. Detecting *that* needs a change-detector rather
|
||||
> than an arrival-detector, and it is not in this version. Prefer `ha_dsmr` where
|
||||
> both are available.
|
||||
|
||||
### Control
|
||||
|
||||
| option | default | meaning |
|
||||
@@ -143,9 +114,9 @@ behind it would trip the firmware watchdog on a system that is working fine.
|
||||
| `step_w` | 10 | Quantisation |
|
||||
| `saturation_w` | 500 | Divergence that counts as "the inverter is at a limit" |
|
||||
| `saturation_cycles` | 3 | How many consecutive cycles before freezing. **Do not set to 1** |
|
||||
| `integrator_max_w` | 0 | Bound on the loop's accumulator, and 0 means "same as `max_w`". Caps how much stale error can be waiting to unwind when the sign flips. **Do not raise it above `max_w`** - the output clamp already bounds what is commanded, so the only thing extra headroom buys is more cycles of wrong-direction power after every saturation event. Lowering it below `max_w` is the useful direction |
|
||||
| `integrator_max_w` | 3000 | Bound on the loop's accumulator, separate from `max_w`. Caps how much stale error can be waiting to unwind when the sign flips. **Keep it above `max_w`, and do not set it equal to `max_w`** |
|
||||
| `heartbeat_s` | 10 | Refresh interval; must stay well under the firmware watchdog |
|
||||
| `stale_input_s` | 15 | How long inputs may be missing before commanding 0 W. In P1 mode this clock starts only *after* `meter_max_age_s` has already expired — the two stack, see "How long a dead meter takes to reach 0 W" |
|
||||
| `stale_input_s` | 15 | How long inputs may be missing before commanding 0 W |
|
||||
| `auto_start` | false | Start controlling on boot (only after commissioning) |
|
||||
|
||||
#### Why `target_grid_w` is not zero
|
||||
|
||||
@@ -28,22 +28,16 @@ class Tuning:
|
||||
step_w: int = 10
|
||||
saturation_w: float = 500.0
|
||||
saturation_cycles: int = 3
|
||||
# The integrator's own bound. None means "follow max_w", which is the
|
||||
# default and the recommended setting.
|
||||
#
|
||||
# ⚠️ DO NOT RAISE THIS ABOVE max_w without a measurement to justify it.
|
||||
# Every watt of integrator above the rail is a watt of wind that has to be
|
||||
# burned off before the command can start moving the other way, i.e. extra
|
||||
# cycles of discharge into an already-exporting meter after every
|
||||
# saturation event. Measured on the closed-loop sim, 4000 W load dropped to
|
||||
# 0: at integrator_max_w == max_w the command is 1000 W two cycles later; at
|
||||
# 1.5x max_w it is 1800 W. The output clamp already bounds what reaches the
|
||||
# wire, so headroom here buys nothing but unwind latency.
|
||||
#
|
||||
# It is a separate key because it has to be able to be SMALLER than max_w,
|
||||
# which is the only direction that buys anything: it caps unwind latency
|
||||
# below what the rail implies. Merging it into max_w would take that away.
|
||||
integrator_max_w: float | None = None
|
||||
# ⚠️ The integrator's OWN bound, and deliberately not max_w. A commercial
|
||||
# controller on this same site clamped only its output and still reported
|
||||
# 14 768 W: with the inverter switched off its integrator climbed ~130 W
|
||||
# every 4 s past 10 kW while the output sat on the 5 kW rail, so the moment
|
||||
# the error flipped there were minutes of accumulated wind to burn off
|
||||
# before the command moved at all. Bounding the accumulator is what makes
|
||||
# recovery time finite; bounding the output only hides it.
|
||||
# Headroom above max_w is wanted (a legitimate large error must not be
|
||||
# truncated at the rail), headroom without limit is the bug.
|
||||
integrator_max_w: float = 3000.0
|
||||
# What the meter should rest at, in W. Negative = a slight export.
|
||||
# ⚠️ The deadband is a one-way ratchet: any resting point inside it holds
|
||||
# forever, and the meter's IMPORT register counts every positive one with
|
||||
@@ -60,9 +54,9 @@ class Decision:
|
||||
sat_count: int
|
||||
frozen: bool
|
||||
reason: str
|
||||
# The integrator AFTER this cycle, before the output clamp, the slew limit
|
||||
# and quantisation. Carry it back in as `i_w` next cycle; that is what keeps
|
||||
# it a separate quantity from the command.
|
||||
# The integrator AFTER this cycle, pre-clamp-to-max_w. Carry it back in as
|
||||
# `i_w` next cycle; that is what keeps it a separate quantity from the
|
||||
# command, which is the whole point of the bound above.
|
||||
i_w: float = 0.0
|
||||
|
||||
|
||||
@@ -95,19 +89,14 @@ def compute(
|
||||
# lets slew be larger than saturation_w.
|
||||
#
|
||||
# The spec states this window twice and differently: "> 10 s" (§11.2) and
|
||||
# "3 samples" (§10.3). This counts CYCLES, and a cycle is not a unit of
|
||||
# time: run_control() calls cycle() only when the meter value CHANGES
|
||||
# (`if self.grid != last_grid`), so three cycles is three distinct meter
|
||||
# readings and nothing more. At the reference P1's ~5 s update rate that is
|
||||
# usually ~15 s, but there is no upper bound on it - a meter that repeats a
|
||||
# value stalls the counter.
|
||||
#
|
||||
# That is a detection-latency limit, not a windup hazard: the same
|
||||
# condition that stalls the counter stalls the whole loop, so nothing
|
||||
# accumulates in the meantime either. If a wall-clock window is ever
|
||||
# required, it belongs in Controller (which has a clock) and not here.
|
||||
# ponytail: this function is worth keeping clockless; the ceiling is that
|
||||
# saturation_cycles cannot express a guaranteed number of seconds.
|
||||
# "3 samples" (§10.3). Cycles are authoritative here because this function
|
||||
# has no clock - it is driven one cycle per meter update by run_control(),
|
||||
# which only calls cycle() when the meter value changes. At the ~5 s
|
||||
# HomeWizard P1 cadence the default 3 cycles is ~15 s, i.e. the stricter
|
||||
# reading of the two. On a faster meter it is not, so saturation_cycles is
|
||||
# configurable and must be raised to keep the window over 10 s.
|
||||
# ponytail: a seconds-based window would mean plumbing wall-clock or dt
|
||||
# into a pure function whose whole value is that it has neither.
|
||||
saturated_now = abs(prev_w - actual_w) > tuning.saturation_w
|
||||
sat_count = min(sat_count + 1, 10) if saturated_now else 0
|
||||
frozen = sat_count >= tuning.saturation_cycles
|
||||
@@ -123,48 +112,32 @@ def compute(
|
||||
|
||||
# --- the integrator ----------------------------------------------------
|
||||
# This loop is in velocity form: the accumulator IS the commanded power, so
|
||||
# "the integrator" and "the output" were one variable and could not be
|
||||
# bounded apart. `i_w` is that accumulator made explicit; main.py carries it
|
||||
# between cycles, which is what turns the two clamps into two limits.
|
||||
#
|
||||
# Passing i_w=None re-seeds it from the last command every cycle. With
|
||||
# integrator_max_w following max_w that reduces this function to the exact
|
||||
# velocity form it replaced, frozen branch included - asserted by an
|
||||
# exhaustive comparison against a transcription of the old law in
|
||||
# test_control.py, not by inspection. Break either the gate or the bound
|
||||
# below and that test is what tells you the equivalence went with it.
|
||||
# for years "the integrator" and "the output" were one variable and could
|
||||
# not be bounded apart. `i_w` is that accumulator made explicit. A caller
|
||||
# that passes nothing gets the old behaviour exactly - seeded from the last
|
||||
# command every cycle - and main.py carries it instead, which is what turns
|
||||
# the two clamps below into two independent limits.
|
||||
if i_w is None:
|
||||
i_w = float(prev_w)
|
||||
limit = tuning.max_w if tuning.integrator_max_w is None else tuning.integrator_max_w
|
||||
|
||||
if abs(error) < tuning.deadband_w:
|
||||
reason = "deadband"
|
||||
else:
|
||||
moved = i_w + tuning.gain * error
|
||||
# ⚠️ Freeze means "may not wind FURTHER in the direction it is already
|
||||
# pushing". It may fall, cross zero, or reverse outright.
|
||||
#
|
||||
# It must NOT be encoded as "only corrections that shrink |i_w|": that
|
||||
# is unsatisfiable for BOTH signs of error whenever the correction is
|
||||
# larger than twice the integrator, i.e. every time the integrator is
|
||||
# near zero. The loop then sits at its last value forever, because what
|
||||
# clears the freeze is the inverter tracking again and not-tracking is
|
||||
# the definition of saturation. Measured on that encoding: 0 W held
|
||||
# indefinitely into a 2 kW import, where this form recovers next cycle.
|
||||
#
|
||||
# This is the same asymmetric rule the output freeze uses below, which
|
||||
# has been in service on real hardware. It is applied here as well
|
||||
# because the requirement is that the INTEGRATOR stop accumulating, not
|
||||
# only the command.
|
||||
if not frozen:
|
||||
i_w = moved
|
||||
else:
|
||||
i_w = min(moved, i_w) if i_w > 0 else max(moved, i_w)
|
||||
step_i = tuning.gain * error
|
||||
# ⚠️ Freeze means "may not wind FURTHER", not "may not move". A strict
|
||||
# freeze would strand the command at whatever it had reached until the
|
||||
# inverter started tracking again - and the inverter is not tracking,
|
||||
# that is what saturation means, so nothing would ever release it. The
|
||||
# unwind direction is the escape route and stays open; the same rule is
|
||||
# applied again to the output below.
|
||||
if not frozen or abs(i_w + step_i) < abs(i_w):
|
||||
i_w = i_w + step_i
|
||||
|
||||
# ⚠️ Applied EVERY cycle, frozen or not: the freeze is conditional, this
|
||||
# bound is not. It is what makes the worst-case unwind time finite and
|
||||
# knowable instead of a function of how long the error happened to stand.
|
||||
i_w = max(-limit, min(limit, i_w))
|
||||
# ⚠️ Applied EVERY cycle, frozen or not, and before the output clamp: the
|
||||
# freeze is conditional, this bound is not. Order matters only in that the
|
||||
# command below is derived from the already-bounded integrator, so no
|
||||
# accumulated value can reach the wire even once.
|
||||
i_w = max(-tuning.integrator_max_w, min(tuning.integrator_max_w, i_w))
|
||||
want = i_w
|
||||
|
||||
# ⚠️ Maintenance shaping (charge-only, cheap-window floor) used to live
|
||||
|
||||
@@ -39,7 +39,7 @@ from .control import Tuning, compute, maintenance_charge_floor, peak_at_risk
|
||||
from .hass import HomeAssistant
|
||||
from .maintenance import IDLE, MaintConfig, Maintenance
|
||||
from .mqtt import MqttPublisher
|
||||
from .p1 import P1Ingest, build_source, is_enabled
|
||||
from .p1 import P1Ingest, build_source
|
||||
from . import web
|
||||
|
||||
OPTIONS_PATH = "/data/options.json"
|
||||
@@ -71,11 +71,7 @@ class Controller:
|
||||
step_w=int(opts.get("step_w", 10)),
|
||||
saturation_w=float(opts.get("saturation_w", 500)),
|
||||
saturation_cycles=int(opts.get("saturation_cycles", 3)),
|
||||
# 0 / unset means "follow max_w", which is the recommended
|
||||
# value. Read the note in control.py before raising it above
|
||||
# max_w: every watt above the rail is unwind latency.
|
||||
integrator_max_w=(float(opts["integrator_max_w"])
|
||||
if opts.get("integrator_max_w") else None),
|
||||
integrator_max_w=float(opts.get("integrator_max_w", 3000)),
|
||||
)
|
||||
self.maint = Maintenance(
|
||||
MaintConfig(
|
||||
@@ -96,7 +92,7 @@ class Controller:
|
||||
# until it opts in.
|
||||
self.p1 = P1Ingest(phases=int(opts.get("meter_phases", 1)),
|
||||
max_age_s=float(opts.get("meter_max_age_s", 30)))
|
||||
self.p1_enabled = is_enabled(opts)
|
||||
self.p1_enabled = str(opts.get("meter_source", "off")) not in ("off", "")
|
||||
|
||||
# live state
|
||||
self.auto = bool(store.data.get("auto", opts.get("auto_start", False)))
|
||||
@@ -324,26 +320,16 @@ class Controller:
|
||||
await asyncio.sleep(1)
|
||||
|
||||
def publish(self) -> None:
|
||||
values = {
|
||||
self.mqtt.publish({
|
||||
"setpoint": self.target,
|
||||
"grid": self.grid,
|
||||
"battery": self.batt,
|
||||
"soc": self.soc,
|
||||
"phase": self.maint.phase,
|
||||
"status": "running" if self.auto else "stopped",
|
||||
}
|
||||
# ⚠️ ONLY when P1 ingestion is actually running. The ESP32's stale-input
|
||||
# watchdog subscribes to sensor.p1_sample_age_s and forces the layer-1
|
||||
# failsafe once it reaches max_age_s. With meter_source off there is no
|
||||
# ingester feeding it, so published_age_s would be time-since-startup
|
||||
# climbing without bound - i.e. every existing install would cross the
|
||||
# threshold within 30 s and pin its inverter at 0 W forever. Publishing
|
||||
# nothing leaves the entity non-existent, which is the status quo and
|
||||
# what has_state() in the firmware is checking for.
|
||||
if self.p1_enabled:
|
||||
# Recomputed here, once a second, on purpose - see P1Ingest.
|
||||
values["p1_age"] = round(self.p1.published_age_s, 1)
|
||||
self.mqtt.publish(values)
|
||||
"p1_age": round(self.p1.published_age_s, 1),
|
||||
})
|
||||
|
||||
async def shutdown(self) -> None:
|
||||
"""Deterministic wind-down. Do not skip this."""
|
||||
@@ -511,7 +497,6 @@ async def amain() -> None:
|
||||
broker.get("port", 1883) if broker else 1883,
|
||||
broker.get("username") if broker else None,
|
||||
broker.get("password") if broker else None,
|
||||
omit=() if is_enabled(opts) else ("p1_age",),
|
||||
)
|
||||
except Exception as err: # noqa: BLE001
|
||||
_LOG.warning("MQTT unavailable (%s) - continuing without status entities", err)
|
||||
|
||||
@@ -60,12 +60,7 @@ AVAILABILITY = f"{BASE}/availability"
|
||||
|
||||
|
||||
class MqttPublisher:
|
||||
def __init__(self, host, port, username=None, password=None, omit=()):
|
||||
# `omit` drops sensor keys from discovery entirely. ⚠️ Announcing a
|
||||
# sensor that nothing will ever publish to is not harmless here:
|
||||
# p1_sample_age_s is a watchdog input, and an entity that exists but is
|
||||
# never fed is a worse signal than one that does not exist at all.
|
||||
self.omit = set(omit)
|
||||
def __init__(self, host, port, username=None, password=None):
|
||||
self.enabled = mqtt is not None and bool(host)
|
||||
self.client = None
|
||||
if not self.enabled:
|
||||
@@ -99,8 +94,6 @@ class MqttPublisher:
|
||||
|
||||
def _announce(self) -> None:
|
||||
for key, object_id, name, unit, dev_class, state_class, icon in SENSORS:
|
||||
if key in self.omit:
|
||||
continue
|
||||
cfg = {
|
||||
"name": name,
|
||||
"object_id": object_id,
|
||||
|
||||
@@ -213,16 +213,8 @@ class QuarterAverager:
|
||||
halves credited to the two blocks, never attributed wholly to either.
|
||||
"""
|
||||
|
||||
def __init__(self, phases: int = 1, max_hold_s: float = 30.0):
|
||||
def __init__(self, phases: int = 1):
|
||||
self.phases = phases
|
||||
# ⚠️ How long one sample may be held forward before the series is
|
||||
# treated as a gap rather than a plateau. Without this the meter can die
|
||||
# while importing 5 kW, come back ten minutes later, and the hold-forward
|
||||
# credits 5 kW x 600 s to the capacity-tariff accumulator - a fabricated
|
||||
# peak, on a permanent record, from data that was never measured. Set
|
||||
# from meter_max_age_s: the point past which the reading is not trusted
|
||||
# for control is the point past which it must not be billed either.
|
||||
self.max_hold_s = float(max_hold_s)
|
||||
self._block: int | None = None # epoch seconds of the block start
|
||||
self._acc = 0.0 # W*s of offtake in the open block
|
||||
self._pp_acc = [0.0] * phases
|
||||
@@ -281,22 +273,16 @@ class QuarterAverager:
|
||||
return closed
|
||||
|
||||
cursor = self._last_t
|
||||
# Beyond this instant the held value stops being evidence of anything.
|
||||
# The stretch from here to `t` is walked so the block boundaries are
|
||||
# still crossed correctly, but nothing is accumulated and `_elapsed`
|
||||
# does not grow - which is what makes a closed block, always divided by
|
||||
# the full 900 s, actually get dragged down by the missing coverage.
|
||||
hold_end = self._last_t + self.max_hold_s
|
||||
while True:
|
||||
end = self._block + QUARTER_S
|
||||
stop = min(t, end)
|
||||
covered = max(0.0, min(stop, hold_end) - cursor)
|
||||
if covered > 0:
|
||||
self._acc += max(self._last_net, 0.0) * covered
|
||||
dt = stop - cursor
|
||||
if dt > 0:
|
||||
self._acc += max(self._last_net, 0.0) * dt
|
||||
if self._last_pp is not None:
|
||||
for i, v in enumerate(self._last_pp[: self.phases]):
|
||||
self._pp_acc[i] += max(v, 0.0) * covered
|
||||
self._elapsed += covered
|
||||
self._pp_acc[i] += max(v, 0.0) * dt
|
||||
self._elapsed += dt
|
||||
cursor = stop
|
||||
if stop < end:
|
||||
break
|
||||
@@ -335,9 +321,7 @@ class P1Ingest:
|
||||
def __init__(self, phases: int = 1, max_age_s: float = 30.0):
|
||||
self.phases = phases
|
||||
self.max_age_s = float(max_age_s)
|
||||
# The same threshold governs control and billing: a reading too old to
|
||||
# steer by is too old to bill by. See QuarterAverager.max_hold_s.
|
||||
self.averager = QuarterAverager(phases, max_hold_s=self.max_age_s)
|
||||
self.averager = QuarterAverager(phases)
|
||||
self.blocks: list[QuarterBlock] = []
|
||||
self.samples = 0
|
||||
self.parse_errors = 0
|
||||
@@ -493,17 +477,9 @@ class HaDsmrSource:
|
||||
continue
|
||||
payload = json.loads(msg.data)
|
||||
if payload.get("id") == 2 and payload.get("type") == "result":
|
||||
# ⚠️ Prime the cache, but do NOT build a sample from it.
|
||||
# get_states returns whatever HA currently holds, which
|
||||
# after a Core restart is a RestoreEntity value of unknown
|
||||
# age. Stamping that with ingest_ts=now resets the age to
|
||||
# zero and reports a fresh meter that may have been dead for
|
||||
# an hour - a synthetic sample hiding the outage from the
|
||||
# watchdog that exists to catch it. The cache is what lets
|
||||
# the FIRST real state_changed build a complete sample; the
|
||||
# age stays honest until one arrives.
|
||||
for obj in payload.get("result") or []:
|
||||
self._absorb(obj.get("entity_id"), obj.get("state"))
|
||||
self._schedule()
|
||||
elif payload.get("type") == "event":
|
||||
data = (payload.get("event") or {}).get("data") or {}
|
||||
if data.get("entity_id") not in self.ids:
|
||||
@@ -671,18 +647,6 @@ class MqttP1Source:
|
||||
# --------------------------------------------------------------------------- #
|
||||
# selection
|
||||
# --------------------------------------------------------------------------- #
|
||||
def is_enabled(opts: dict) -> bool:
|
||||
"""Whether P1 ingestion is switched on at all.
|
||||
|
||||
⚠️ One definition, because three places depend on it and they MUST agree:
|
||||
where the grid reading comes from, whether the ingest task is started, and
|
||||
whether sensor.p1_sample_age_s is announced over MQTT discovery. An age
|
||||
sensor announced with no ingester behind it is a watchdog input nobody is
|
||||
feeding, and the ESP32 trips on it.
|
||||
"""
|
||||
return str(opts.get("meter_source", "off") or "off").strip() not in ("off", "")
|
||||
|
||||
|
||||
def build_source(opts: dict, ingest: P1Ingest, session, broker: dict | None):
|
||||
"""Return the transport named by `meter_source`, or None if disabled.
|
||||
|
||||
@@ -690,7 +654,7 @@ def build_source(opts: dict, ingest: P1Ingest, session, broker: dict | None):
|
||||
changing transport is a config edit, never a code path.
|
||||
"""
|
||||
source = str(opts.get("meter_source", "off") or "off").strip()
|
||||
if not is_enabled(opts):
|
||||
if source in ("off", ""):
|
||||
return None
|
||||
if source == SOURCE_HA:
|
||||
return HaDsmrSource(session, ingest, {
|
||||
|
||||
@@ -65,7 +65,7 @@ options:
|
||||
step_w: 10
|
||||
saturation_w: 500
|
||||
saturation_cycles: 3
|
||||
integrator_max_w: 0
|
||||
integrator_max_w: 3000
|
||||
heartbeat_s: 10
|
||||
stale_input_s: 15
|
||||
auto_start: false
|
||||
@@ -121,7 +121,7 @@ schema:
|
||||
step_w: int(1,100)
|
||||
saturation_w: int(100,2000)
|
||||
saturation_cycles: int(1,10)
|
||||
integrator_max_w: int(0,15000)
|
||||
integrator_max_w: int(100,15000)
|
||||
heartbeat_s: int(2,25)
|
||||
stale_input_s: int(5,120)
|
||||
auto_start: bool
|
||||
|
||||
@@ -78,7 +78,7 @@ print("SAFETY-04: the integrator is bounded apart from the output")
|
||||
# The historical runaway, with its real numbers. A commercial controller on
|
||||
# this site, with the inverter switched OFF, wound ~130 W every 4 s past 10 kW
|
||||
# and reported 14 768 W while its output clamp sat at 5 kW. At gain 0.6 that
|
||||
# rate is a standing error of 130/0.6 = 217 W that never resolves, because the
|
||||
# rate is a standing error of 130/0.6 ≈ 217 W that never resolves, because the
|
||||
# inverter is not there to resolve it. 150 cycles is past the ~113 it took to
|
||||
# reach 14 768 W at that rate.
|
||||
RUNAWAY_ERROR = 130.0 / 0.6
|
||||
@@ -99,53 +99,34 @@ def runaway(tuning):
|
||||
return worst_i, worst_cmd
|
||||
|
||||
|
||||
TR = Tuning(max_w=2000) # integrator_max_w unset => follows max_w
|
||||
TR = Tuning(max_w=2000, integrator_max_w=3000)
|
||||
wi, wc = runaway(TR)
|
||||
check(f"runaway: integrator plateaus at {wi:.0f} W (<= 2000)", wi <= TR.max_w)
|
||||
check(f"runaway: integrator plateaus at {wi:.0f} W (<= 3000)", wi <= TR.integrator_max_w)
|
||||
check(f"runaway: emitted command peaks at {wc:.0f} W (<= 2000)", wc <= TR.max_w)
|
||||
check("runaway: nowhere near the historical 14 768 W", wc < HISTORICAL_W / 4)
|
||||
|
||||
# ...and with the saturation detector deliberately defeated, so that only the
|
||||
# clamp is holding. Kill one mechanism, the other still bounds it.
|
||||
TD = Tuning(max_w=2000, saturation_w=1e9)
|
||||
# clamp is holding. This is the AC that says the two mechanisms are
|
||||
# independent: kill one, the other still bounds it.
|
||||
TD = Tuning(max_w=2000, integrator_max_w=3000, saturation_w=1e9)
|
||||
wi, wc = runaway(TD)
|
||||
check(f"runaway with the detector defeated: integrator still bounded ({wi:.0f} W)",
|
||||
wi <= TD.max_w)
|
||||
check(f"runaway with the detector defeated: integrator still <= 3000 ({wi:.0f} W)",
|
||||
wi <= TD.integrator_max_w)
|
||||
check("runaway with the detector defeated: command still <= max_w", wc <= TD.max_w)
|
||||
|
||||
# The bound is a separate quantity, and the useful direction is BELOW max_w:
|
||||
# there it binds first and caps unwind latency tighter than the rail does.
|
||||
# The bound is not max_w. If someone "simplifies" them into one key this fails.
|
||||
d = compute(prev_w=0, grid_w=6000, actual_w=0,
|
||||
tuning=Tuning(max_w=2000, integrator_max_w=1000, slew_w=5000))
|
||||
check("integrator bound binds independently of the output clamp",
|
||||
d.i_w == 1000 and d.target_w == 1000)
|
||||
tuning=Tuning(max_w=2000, integrator_max_w=3000, slew_w=5000))
|
||||
check("integrator bound is separate from the output clamp",
|
||||
d.i_w == 3000 and d.target_w == 2000)
|
||||
|
||||
# Freeze = may not wind further in the direction it is already pushing.
|
||||
# Freeze = does not accumulate. Same input twice; the integrator must not move.
|
||||
TF = Tuning(saturation_w=500, saturation_cycles=3)
|
||||
f1 = compute(prev_w=2000, grid_w=800, actual_w=0, tuning=TF, sat_count=3, i_w=2000.0)
|
||||
check("frozen: integration does not wind further", f1.i_w == 2000.0 and f1.frozen)
|
||||
check("frozen: integration does not accumulate", f1.i_w == 2000.0 and f1.frozen)
|
||||
f2 = compute(prev_w=2000, grid_w=-800, actual_w=0, tuning=TF, sat_count=3, i_w=2000.0)
|
||||
check("frozen: unwinding is still allowed", f2.i_w < 2000.0)
|
||||
|
||||
# ⚠️ REGRESSION, and the reason the first cut of SAFETY-04 was rejected. A
|
||||
# freeze encoded as "only corrections that shrink |i_w|" is unsatisfiable for
|
||||
# BOTH signs of error whenever |correction| > 2*|i_w|, so near zero the loop
|
||||
# stops moving forever - the freeze cannot clear, because clearing it needs the
|
||||
# inverter to track and not-tracking is what saturation means. Measured on that
|
||||
# encoding: 0 W held into a 2 kW import for as long as the sim ran.
|
||||
z = compute(prev_w=0, grid_w=2000, actual_w=600, tuning=T, sat_count=3, i_w=0.0)
|
||||
check("frozen at i_w=0: a 2 kW import still moves the command",
|
||||
z.frozen and z.target_w == 1000)
|
||||
# ...and the next cycle the inverter is inside saturation_w of the command, so
|
||||
# the freeze clears on its own. Deadlock would show up here as frozen=True.
|
||||
z2 = compute(prev_w=1000, grid_w=1000, actual_w=600, tuning=T,
|
||||
sat_count=z.sat_count, i_w=z.i_w)
|
||||
check("frozen at i_w=0: the freeze then clears", not z2.frozen)
|
||||
# Same stranding on the other side: a small positive integrator against export.
|
||||
z3 = compute(prev_w=100, grid_w=-1000, actual_w=800, tuning=T, sat_count=3, i_w=100.0)
|
||||
check("frozen at i_w=+100: a 1 kW export still moves the command",
|
||||
z3.frozen and z3.target_w < 0)
|
||||
|
||||
# False-positive guard: a normal 2 kW load step must not trip the detector,
|
||||
# because the plant needs several cycles to catch up on every one of them.
|
||||
prev, actual, sat, i_w, froze = 0.0, 0.0, 0, 0.0, False
|
||||
@@ -156,74 +137,6 @@ for _ in range(12):
|
||||
froze = froze or d.frozen
|
||||
check("a normal 2 kW load step does not trip the saturation freeze", not froze)
|
||||
|
||||
# The convergence sim below runs WITHOUT a carried integrator. This is the same
|
||||
# 2 kW step in the configuration that actually ships, where main.py carries it.
|
||||
prev, actual, sat, i_w = 0.0, 0.0, 0, 0.0
|
||||
carried = 0
|
||||
for _ in range(12):
|
||||
d = compute(prev, 2000.0 - actual, actual, T, sat, i_w)
|
||||
prev, sat, i_w = d.target_w, d.sat_count, d.i_w
|
||||
actual = actual + 0.94 * (prev - actual)
|
||||
carried += 1
|
||||
if abs(2000.0 - actual) < T.deadband_w:
|
||||
break
|
||||
check(f"carried integrator converges in {carried} cycles (<=6)", carried <= 6)
|
||||
check("carried integrator does not overshoot the load", actual <= 2000.0 + T.deadband_w)
|
||||
|
||||
# ⚠️ REGRESSION: an integrator allowed to wind past the rail buys nothing (the
|
||||
# output clamp already bounds the wire) and costs extra cycles of
|
||||
# wrong-direction power after every saturation event. 4000 W load held to
|
||||
# saturation, then dropped to 0; the figure is the command on the first cycle
|
||||
# after the drop. This is what makes the DOCS advice checkable.
|
||||
def unwind(t):
|
||||
prev, actual, sat, i_w, load = 0.0, 0.0, 0, 0.0, 4000.0
|
||||
for c in range(16):
|
||||
if c == 15:
|
||||
load = 0.0
|
||||
d = compute(prev, load - actual, actual, t, sat, i_w)
|
||||
prev, sat, i_w = d.target_w, d.sat_count, d.i_w
|
||||
actual = actual + 0.94 * (prev - actual)
|
||||
return prev
|
||||
|
||||
|
||||
tight, loose = unwind(Tuning(max_w=2000)), unwind(Tuning(max_w=2000, integrator_max_w=3000))
|
||||
check(f"after saturation ends the command is {tight:.0f} W (<= 1000)", tight <= 1000)
|
||||
check(f"headroom above max_w makes that worse ({loose:.0f} W) - hence the default",
|
||||
loose > tight)
|
||||
|
||||
print("SAFETY-04: the i_w=None path is still release/1.0, exactly")
|
||||
|
||||
|
||||
def legacy(prev, grid, actual, t, sat_count):
|
||||
"""release/1.0's control law, transcribed. Do not 'improve' this."""
|
||||
sc = min(sat_count + 1, 10) if abs(prev - actual) > t.saturation_w else 0
|
||||
frozen = sc >= t.saturation_cycles
|
||||
error = grid - t.target_grid_w
|
||||
want = prev if abs(error) < t.deadband_w else prev + t.gain * error
|
||||
target = max(-t.max_w, min(t.max_w, want))
|
||||
target = max(prev - t.slew_w, min(prev + t.slew_w, target))
|
||||
if frozen:
|
||||
target = min(target, prev) if prev > 0 else max(target, prev)
|
||||
step = max(1, int(t.step_w))
|
||||
return float(round(target / step) * step), sc
|
||||
|
||||
|
||||
# Exhaustive over the interesting corners, both freeze states, both signs, and
|
||||
# either side of the deadband. This is what makes the claim in control.py's
|
||||
# integrator comment a checked fact rather than an assertion.
|
||||
diffs = []
|
||||
for tune in (Tuning(), Tuning(target_grid_w=-10.0), Tuning(max_w=5000, slew_w=5000)):
|
||||
for prev in (-2000.0, -500.0, -100.0, 0.0, 100.0, 500.0, 2000.0):
|
||||
for grid in (-6000.0, -1000.0, -500.0, -14.0, 0.0, 14.0, 500.0, 1000.0, 6000.0):
|
||||
for actual in (-2000.0, 0.0, 600.0, 2000.0):
|
||||
for sc in (0, 2, 3, 9):
|
||||
d = compute(prev, grid, actual, tune, sc) # i_w defaults to None
|
||||
lt, lsc = legacy(prev, grid, actual, tune, sc)
|
||||
if (d.target_w, d.sat_count) != (lt, lsc):
|
||||
diffs.append((prev, grid, actual, sc, d.target_w, lt))
|
||||
check(f"i_w=None reproduces release/1.0 over {3*7*9*4*4} cases"
|
||||
+ (f" (first diff {diffs[0]})" if diffs else ""), not diffs)
|
||||
|
||||
print("capacity tariff")
|
||||
check("no forecast means no cap", maintenance_charge_floor(2500, None, 3500) == 2500)
|
||||
check("headroom caps the charge", maintenance_charge_floor(2500, 2000, 3500) == 1500)
|
||||
|
||||
@@ -229,47 +229,6 @@ before = a.partial_ws
|
||||
a.add(sample(1000.0, at=BASE + timedelta(seconds=5)))
|
||||
check("an out-of-order telegram is dropped, not integrated backwards",
|
||||
a.partial_ws == before and a.elapsed_s == 10.0)
|
||||
# ⚠️ The assertion above is NOT sufficient on its own, and that is the whole
|
||||
# lesson: deleting the guard still passes it, because the negative interval is
|
||||
# separately refused by the `covered > 0` test. What the guard actually prevents
|
||||
# is the REWIND - without it the held timestamp moves back to +5 s and the next
|
||||
# telegram re-integrates the 5..10 s window that was already counted. The damage
|
||||
# only becomes visible one sample later, so the test has to go one sample later.
|
||||
a.add(sample(1000.0, at=BASE + timedelta(seconds=20)))
|
||||
check("...and the held timestamp is not rewound, so the next telegram "
|
||||
"cannot double-count", a.elapsed_s == 20.0 and a.partial_ws == 20000.0)
|
||||
|
||||
# A duplicate telegram (identical timestamp) is the same rule.
|
||||
a = QuarterAverager(1)
|
||||
a.add(sample(1000.0, at=BASE))
|
||||
a.add(sample(1000.0, at=BASE + timedelta(seconds=10)))
|
||||
a.add(sample(4000.0, at=BASE + timedelta(seconds=10)))
|
||||
a.add(sample(1000.0, at=BASE + timedelta(seconds=20)))
|
||||
check("a duplicate timestamp neither re-integrates nor replaces the held value",
|
||||
a.elapsed_s == 20.0 and a.partial_ws == 20000.0)
|
||||
|
||||
# A gap must not be filled with the last held value. The meter dies at 5 kW and
|
||||
# returns ten minutes later; hold-forward would credit 5 kW x 600 s to the
|
||||
# capacity-tariff accumulator - a fabricated peak, on a permanent record, from
|
||||
# data nobody measured.
|
||||
a = QuarterAverager(1, max_hold_s=30.0)
|
||||
a.add(sample(5000.0, at=BASE))
|
||||
a.add(sample(5000.0, at=BASE + timedelta(seconds=600)))
|
||||
check("a 600 s gap is held for at most max_hold_s, not for the whole gap",
|
||||
a.partial_ws == 5000.0 * 30.0)
|
||||
check("the unobserved stretch does not count as elapsed time", a.elapsed_s == 30.0)
|
||||
closed = a.add(sample(5000.0, at=BASE + timedelta(seconds=900)))
|
||||
check("the outage drags the billed quarter down instead of inventing a peak",
|
||||
len(closed) == 1 and abs(closed[0].offtake_avg_w - 300000.0 / 900.0) < 1e-9)
|
||||
check("...nowhere near the 5000 W a hold-forward would have billed",
|
||||
closed[0].offtake_avg_w < 400.0)
|
||||
|
||||
# The cap must not disturb a normally-spaced stream.
|
||||
a = QuarterAverager(1, max_hold_s=30.0)
|
||||
for i in range(0, 121, 5): # a healthy 5 s telegram cadence
|
||||
a.add(sample(2000.0, at=BASE + timedelta(seconds=i)))
|
||||
check("a healthy 5 s cadence is untouched by the hold cap",
|
||||
a.elapsed_s == 120.0 and abs(a.offtake_avg_w - 2000.0) < 1e-9)
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
print("ingest timestamp, age and staleness")
|
||||
@@ -523,15 +482,9 @@ async def _e2e():
|
||||
|
||||
live, wire = asyncio.run(_e2e())
|
||||
check("the websocket handshake and subscription complete", live.samples >= 1)
|
||||
# ⚠️ TWO, not three. get_states primes the cache but must NOT build a sample:
|
||||
# HA returns whatever it currently holds, which after a Core restart is a
|
||||
# RestoreEntity value of unknown age, and stamping that with ingest_ts=now
|
||||
# resets the age and reports a fresh meter that may have been dead for an hour.
|
||||
# Only the two real state_changed telegrams become samples. Four state_changed
|
||||
# events arrived (two per telegram); the debounce is what makes those two
|
||||
# consistent samples rather than four half-updated ones.
|
||||
check("connecting does not manufacture a sample from cached HA state",
|
||||
live.samples == 2)
|
||||
# Six state_changed events arrived (two per telegram). The debounce is what
|
||||
# makes that three consistent samples instead of six half-updated ones.
|
||||
check("three telegrams produce three samples, not six", live.samples == 3)
|
||||
check("the final export-dominant telegram nets negative",
|
||||
live.last.net_w == -800.0)
|
||||
check("the sample was built over the wire, tagged with its transport",
|
||||
@@ -539,102 +492,9 @@ check("the sample was built over the wire, tagged with its transport",
|
||||
check("an entity we did not subscribe to is never cached",
|
||||
"sensor.something_else" not in wire.cache and len(wire.cache) == 1)
|
||||
check("a mid-stream unavailable is a parse error, not a sample",
|
||||
live.parse_errors == 1 and live.samples == 2)
|
||||
live.parse_errors == 1 and live.samples == 3)
|
||||
check("the last good reading survives the unavailable", live.net_w == -800.0)
|
||||
check("the averager integrated the live stream", live.averager.elapsed_s > 0.2)
|
||||
|
||||
# The reason get_states still matters: it is what lets the FIRST real telegram
|
||||
# build a complete sample instead of waiting for every entity to change once.
|
||||
check("the primed cache let the first telegram build immediately",
|
||||
live.samples == 2 and live.last.import_w == 0.0)
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
print("the age sensor must not exist when P1 is off")
|
||||
# ⚠️ This is a fleet-wide regression guard, not a nicety. The ESP32 watchdog
|
||||
# does `id(p1_age_s).has_state() && id(p1_age_s).state >= max_age_s` and forces
|
||||
# the layer-1 failsafe. published_age_s counts from P1Ingest.__init__, so if the
|
||||
# age were published with meter_source off it would climb past 30 s on every
|
||||
# existing install within half a minute and pin the inverter at 0 W forever.
|
||||
|
||||
from app.p1 import is_enabled # noqa: E402
|
||||
from app.mqtt import SENSORS, MqttPublisher # noqa: E402
|
||||
|
||||
check("meter_source off is disabled", is_enabled({"meter_source": "off"}) is False)
|
||||
check("a missing meter_source is disabled", is_enabled({}) is False)
|
||||
check("an empty meter_source is disabled", is_enabled({"meter_source": ""}) is False)
|
||||
check("ha_dsmr is enabled", is_enabled({"meter_source": SOURCE_HA}) is True)
|
||||
check("mqtt_p1 is enabled", is_enabled({"meter_source": SOURCE_MQTT}) is True)
|
||||
|
||||
# The entity id SAFETY-01's firmware subscribes to, pinned by object_id.
|
||||
row = [s for s in SENSORS if s[0] == "p1_age"]
|
||||
check("the age sensor is declared exactly once", len(row) == 1)
|
||||
check("its object_id pins entity_id to sensor.p1_sample_age_s",
|
||||
row[0][1] == "p1_sample_age_s")
|
||||
check("it is published in seconds", row[0][3] == "s")
|
||||
|
||||
|
||||
class _RecordingClient:
|
||||
def __init__(self):
|
||||
self.sent = []
|
||||
|
||||
def publish(self, topic, payload=None, retain=False):
|
||||
# Topic AND payload: object_id, the thing that actually pins the entity
|
||||
# id, only appears in the discovery payload. Recording topics alone made
|
||||
# the "is not announced" check pass for the wrong reason.
|
||||
self.sent.append(f"{topic} {payload}")
|
||||
|
||||
|
||||
def _announced(omit):
|
||||
pub = MqttPublisher(None, 1883, omit=omit) # host None -> never connects
|
||||
pub.client = _RecordingClient()
|
||||
pub._announce()
|
||||
return " ".join(pub.client.sent)
|
||||
|
||||
|
||||
check("with P1 off the age sensor is never announced",
|
||||
"p1_sample_age_s" not in _announced(("p1_age",)))
|
||||
check("the other status entities are still announced with P1 off",
|
||||
"goodwe_grid_power" in _announced(("p1_age",)))
|
||||
check("with P1 on the age sensor IS announced",
|
||||
"p1_sample_age_s" in _announced(()))
|
||||
|
||||
# And the publish dict itself, through the real Controller.
|
||||
from app.main import Controller # noqa: E402
|
||||
|
||||
|
||||
class _Store:
|
||||
data = {}
|
||||
|
||||
def set(self, *a):
|
||||
pass
|
||||
|
||||
def get_time(self, *a):
|
||||
return None
|
||||
|
||||
|
||||
class _Pub:
|
||||
def __init__(self):
|
||||
self.last = {}
|
||||
|
||||
def publish(self, values):
|
||||
self.last = values
|
||||
|
||||
def close(self):
|
||||
pass
|
||||
|
||||
|
||||
pub_off = _Pub()
|
||||
Controller({"meter_source": "off"}, None, _Store(), pub_off).publish()
|
||||
check("with P1 off, p1_age is absent from the published payload",
|
||||
"p1_age" not in pub_off.last)
|
||||
check("...while the normal status keys are still published",
|
||||
"setpoint" in pub_off.last and "grid" in pub_off.last)
|
||||
|
||||
pub_on = _Pub()
|
||||
Controller({"meter_source": SOURCE_HA}, None, _Store(), pub_on).publish()
|
||||
check("with P1 on, p1_age is published", "p1_age" in pub_on.last)
|
||||
check("...as a number, so has_state() becomes true only once we feed it",
|
||||
isinstance(pub_on.last["p1_age"], float))
|
||||
check("the averager integrated the live stream", live.averager.elapsed_s > 0.5)
|
||||
|
||||
print()
|
||||
if fails:
|
||||
|
||||
Reference in New Issue
Block a user