Files
goodwe-addon/goodwe_controller/CHANGELOG.md
T
glenn schrooyenandClaude Opus 5 e663e10245 TEL-04 review: unshadow the helper, validate config at startup, and record
what the rig proved about the age sensor

Review findings 1, 3, 5 and 6. Finding 2 is deliberately untouched - it is
its own ticket.

3. `built` was rebound at test_p1.py:816 by `built = build_source(...)`,
   silently disarming the build() wrapper for anything appended below it.
   Renamed to `sel`. Reproduced the reviewer's failure before fixing:
   appending a check that calls built() after that line gives
   `TypeError: 'HaSignedSource' object is not callable` and aborts at 163 of
   180; with the rename the same probe reaches 180 and passes.

5. DOCS.md now states the "length must equal meter_phases" constraint that
   config.yaml already carried, plus what leaving the list empty actually
   costs: on the surveyed reading the phases carry 2769 W of import while the
   connection nets 187 W, so the tariff quantity is understated ~15x.

6. build_source now checks the ha_signed wiring once at startup instead of
   once per telegram: a blank p1_net_entity, or a phase list whose length
   disagrees with meter_phases, logs an error and disables ingestion. Both
   otherwise fail in the single way indistinguishable from a healthy source
   nobody has fed yet - no samples, a climbing age, the watchdog holding the
   battery at 0 W, and nothing in the log.

1. THE AGE SENSOR. Measured on the ENV-01 rig against the real HomeWizard
   integration, meter frozen via hwsim's `?fault=freeze` seam (cleared in a
   finally:, rig verified restored):

     - websocket state_changed for the meter over 70 s : 0
     - last_reported advanced (REST serialiser)        : no
     - last_reported advanced (websocket serialiser)   : no
     - subscribe_events(state_reported)                : rejected,
       "Event filter is required for event state_reported"

   So Home Assistant exposes NO arrival signal for a repeated reading, and
   the proposed fix - stamp from last_reported via subscribe_entities - is
   not available. subscribe_entities listens only to EVENT_STATE_CHANGED, and
   as_compressed_state carries no last_reported at all.

   The age is therefore "time since the value changed", which on ha_dsmr is
   mostly harmless (a telegram moves several entities) and on ha_signed is
   not: one entity means a healthy meter under a flat load is
   indistinguishable from a dead one. Recorded loudly in DOCS.md, in the
   HaSignedSource docstring and in the CHANGELOG, with the measured 42.2 s
   and 97.0 s gaps from our own capture.

   meter_max_age_s is deliberately NOT widened. The two conditions produce an
   identical signal, so a larger number does not separate them - it only
   chooses which of the two errors you get, and it would disarm the watchdog
   for a genuinely dead meter as well. The honest fix is an arrival stamp the
   meter itself provides.

test_p1.py: 174 -> 179 checks, all green. Other three suites unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Du77usMj8XNKNFZGmUiWDa
2026-08-25 17:34:37 +02:00

174 lines
9.4 KiB
Markdown

# Changelog
## Unreleased
**TEL-04.** A third `meter_source`, `ha_signed`, reading **one signed** Home
Assistant entity: positive = import, negative = export. That is the shape a
HomeWizard P1 publishes (`sensor.p1_meter_active_power`), and it is the meter
actually fitted here - which neither TEL-01 transport can read, because
`ha_dsmr` needs two unsigned registers and refuses a negative one, i.e. every
exporting telegram. Set `p1_net_entity`, and `p1_phase_net_entities` for the
per-phase capacity-tariff figures on a three-phase connection.
Everything TEL-01 established is inherited rather than re-implemented - the
new transport is a subclass of the `ha_dsmr` one overriding only which
entities it wants and how they become a sample. So ingest timestamping,
`meter_max_age_s`, `sensor.p1_sample_age_s` recomputed against the clock and
republished once a second, the plausibility bounds, and `unavailable` /
`unknown` treated as a *missing reading and never 0 W* all behave identically
across the three sources.
Still defaults to `off`; an existing install is unaffected until it opts in.
⚠️ **`sensor.p1_sample_age_s` is published on `ha_signed`, but must not yet be
thresholded by the ESP32 stale-input watchdog.** On the HA WebSocket paths the
age is stamped from `state_changed`, so it measures time since the value
*changed*, not since the meter *reported* - and Home Assistant exposes no
arrival signal for a repeated reading (no `state_changed`, no `last_reported`
movement on either serialiser, and `state_reported` is not subscribable over
the WebSocket). Measured on the ENV-01 rig against the real HomeWizard
integration. `ha_dsmr` mostly escapes it because a telegram moves several
entities at once; `ha_signed` has one, so a healthy meter under a flat load is
indistinguishable from a dead one. Our own capture has the house meter going
42.2 s and 97.0 s between changes. Raising `meter_max_age_s` does not fix that,
it only chooses which error you get; the fix is an arrival stamp from the meter
itself and is a separate ticket. Full detail in DOCS.md.
## 0.3.0
**SAFETY-04.** The control law's integrator is now an explicit accumulator,
bounded independently of the output clamp instead of inheriting whatever
headroom the clamp happened to leave. It also freezes while the inverter is
not tracking, rather than continuing to wind up against a command nothing is
acting on. `integrator_max_w` (default `0`) governs the bound; `0` means
"follow `max_w`", which is the existing behaviour.
Behaviour is unchanged at the defaults - a 4,928-case equivalence sweep
against the previous control law confirms it decides identically at
`integrator_max_w: 0`.
**TEL-01.** P1 meter ingestion, so a Belgian P1's two unsigned registers
(consumption, injection) no longer need a hand-written signed template
sensor: the subtraction moves into the add-on, done once and tested. Two
transports, chosen with the new `meter_source` option: `ha_dsmr` subscribes
to the DSMR integration over the HA WebSocket, `mqtt_p1` reads a topic.
Defaults to `off`, which keeps the existing `meter_entity` path untouched -
nothing changes for an install that does not opt in.
Enabling it publishes `sensor.p1_sample_age_s`: seconds since the newest
accepted telegram, recomputed against the clock and republished roughly once
a second rather than only when a telegram lands. That is deliberate - Home
Assistant only pushes a state on change, so a meter sitting at a genuinely
constant reading would otherwise look identical to a dead one. Watching the
age instead means a frozen meter shows a climbing age, not a flat line. The
firmware watchdog subscribes to this exact entity id.
Known limits, both already in DOCS.md: on `mqtt_p1`, a bridge stuck
republishing its last telegram still "arrives", so the age cannot detect
that particular failure - prefer `ha_dsmr` where both are available. And a
dead P1 meter takes 45 s to reach 0 W commanded (30 s for `meter_max_age_s`
to call the reading stale, then 15 s of `stale_input_s` on top), which is
`meter_max_age_s` and `stale_input_s` stacking, not either one alone.
## 0.2.1
`target_grid_w` (default -10 W): what the meter should rest at. The deadband
holds any resting point inside it indefinitely, and the meter bills import and
export on separate registers, so resting at +14 W import costs 0.34 kWh/day
with the loop behaving perfectly. Biasing the target slightly negative moves
that residue onto the export register. Configurable in the Configuration tab;
see DOCS.md for the trade-off table.
Behaviour is unchanged at `target_grid_w: 0`.
## 0.2.0
Precedence between strategies is now a first-class object instead of an if/else
ladder, ahead of there being more than three of them.
Every strategy returns a CLAIM each cycle - `set` ("I want X") or `limit` ("the
result must stay within these bounds") - and `arbiter.py` resolves them by one
rule:
1. Highest-priority `set` wins; no claim at all means 0 W.
2. Then every `limit` whose priority is >= that set's priority applies, most
restrictive first.
3. Contradictory limits are a BUG: command 0 W and say so.
Clause 2 is why "money outranks maintenance" is now a consequence of the
priorities rather than a special case in a Jinja template: the maintenance
charge-only limit binds the loop, but will not bind a higher-priority peak
shaving claim when one exists.
- Maintenance shaping (charge-only, cheap-window floor) moved out of the control
law. `control.py` is once again only a controller that tracks the meter.
- The loop computes from the ARBITER's last output, not its own last wish. If
something outranked it, that is what the hardware actually did, and tracking
anything else makes it jump when it regains control.
- Every decision is explainable: "loop -> 0 W, limited by maintenance
(charge-only)" now appears in the UI and the log, instead of a bare number.
- Safety limits (device rating, supervised max_w) bind every strategy including
the highest, and are still enforced a second time at the point of writing.
## 0.1.6
Findings from installing this on a live system, replacing a working YAML
implementation. Every one of these was silent - the add-on looked healthy while
being completely unable to do its job.
- **`run.sh` must use `#!/usr/bin/with-contenv sh`.** The HA base images run
s6-overlay, which starts services with a SANITISED environment. With a plain
shebang, SUPERVISOR_TOKEN is simply absent and every Home Assistant call
returns 401 - while `homeassistant_api: true` makes the permissions look
correctly granted. Startup now logs the token length and probes the Core API,
so the next person sees it in one line.
- **Bump `version:` for every change.** Supervisor keys the built image by
version, so editing source and rebuilding silently reuses the old image. Two
fixes appeared not to work because of this.
- **Dependencies come from apk, not pip.** Alpine is musl and there are no musl
wheels for aiohttp; pip would compile it on the client's Pi.
- **paho-mqtt 1.x and 2.x are both supported.** Alpine ships 1.x, which has no
`CallbackAPIVersion`; that raised and took the whole add-on down with it.
- **MQTT can no longer take down control.** Publisher construction is wrapped -
observability must never stop the controller.
- **MQTT discovery is published from `on_connect`.** paho drops QoS-0 publishes
issued before the CONNACK, so announcing straight after `connect()` published
nothing at all while logging "MQTT connected".
- **Repeated failures log at most once a minute.** The control loop retries every
second; unthrottled warnings rolled the log buffer and destroyed the startup
diagnostics needed to debug the 401 above.
- **`auto_start` works.** The store's defaults supplied `auto: False`, so the
fallback to the option could never fire.
Known issue: after deleting the MQTT entities from the registry during
development, Home Assistant would not re-adopt them from retained discovery -
not even after clearing the retained topics and reconnecting. The add-on
publishes correct discovery and live state (verified on the broker); this is an
HA-side adoption problem and affects status entities only, never control.
## 0.1.0
First packaged release. Ports the control loop and the monthly maintenance
cycle from the reference Home Assistant implementation into an add-on.
- Grid-following control: gain/slew/clamp/deadband with anti-windup, all tuned
against measured hardware behaviour (see FIELD-GUIDE.md §14).
- Saturation freeze **with the duration term** — three consecutive diverging
cycles, not one. The instantaneous test fires on every large correction,
because the plant itself needs 3-6 s to settle.
- Monthly maintenance cycle as an ownership state machine: drain / charge /
hold, with exactly one writer of the setpoint at any moment.
- Capacity-tariff awareness: the maintenance charge is capped by quarter-hour
peak headroom, and peak shaving outranks the maintenance schedule.
- Failsafe behaviour: commands 0 W on missing inputs, on stop, and on shutdown.
Never replays a stale setpoint - the reference implementation did, and the
hardware watchdog cannot catch that.
- Ingress UI with a commissioning checklist that names problems in words.
- Optional MQTT discovery for status entities.
Known limits:
- Home Assistant OS / Supervised only (add-ons cannot run on Container/Core).
- The inverter protocol is reverse-engineered; no vendor contract.
- Without the optional RS485 e-stop, nothing covers the host machine dying.