Files
goodwe-addon/goodwe_controller/DOCS.md
T
glenn schrooyenandClaude Opus 5 8b51a51e20 TEL-04: a third meter_source for a single signed entity
TEL-01 shipped ha_dsmr and mqtt_p1, and neither can read the meter that is
actually fitted here. The house has a HomeWizard P1 exposing ONE signed
entity, sensor.p1_meter_active_power (+ import, - export); ha_dsmr wants two
unsigned registers and refuses a negative one outright, which is every
exporting telegram. So sensor.p1_sample_age_s could not be produced at this
site, and FW-01's watchdog needs it - measured, not theoretical: the house P1
went 51.1 s and 36.2 s without a state change overnight, both past
meter_max_age_s 30, so without the age sensor the watchdog would false-trip
the battery to 0 W.

Adds meter_source: ha_signed, reading p1_net_entity (and optionally
p1_phase_net_entities in L1..L3 order for the capacity-tariff peak). The
derivation is split_signed(), sitting next to make_sample's subtraction for
the same reason it does - the moment a user is asked to write two template
sensors that split a signed value, the sign convention is back in unreviewed
YAML underneath a safety input, which is exactly what TEL-01 removed.

The transport is a subclass of HaDsmrSource overriding only _wanted() and
build(), so every rule TEL-01 established is inherited rather than
re-implemented: ingest timestamping, meter_max_age_s, the clock-recomputed
sensor.p1_sample_age_s republished ~1 Hz, the plausibility ceiling, the
"prime the cache from get_states but never build a sample out of it" rule,
"a reconnect emits nothing", and unavailable/unknown treated as a MISSING
reading and never as 0 W.

Defaults to off. An existing install is unaffected until it opts in.

test_p1.py: 122 -> 174 checks. Includes an end-to-end run of the new
transport against a fake Home Assistant websocket, and the sign convention
asserted against real captured readings from
sim/scenarios/ha-p1_meter_active_power-2026-08-{20,23}.json (-5710 W at
13:46 local under full sun is export; +775 W at midnight is import).

Non-vacuity: ten mutations of the new rules, each applied alone and reverted
byte-identical. Nine turn the suite red. The tenth - splitting the per-phase
signed values rather than passing them through - is an equivalent mutant,
because make_sample subtracts the two lists again and does not sign-check
per-phase figures. That is recorded in a ponytail: comment at the site rather
than left for the next reviewer to rediscover.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Du77usMj8XNKNFZGmUiWDa
2026-08-25 16:44:22 +02:00

260 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# GoodWe RS485 Controller
Drives a GoodWe ES/BP battery inverter over its RS485 meter bus: holds net grid
exchange at zero, and runs a monthly battery maintenance cycle so the BMS can
balance cells and recalibrate its coulomb counter.
Installers: read `FIELD-GUIDE.md` in the repository. It is not optional reading —
it contains the commissioning gates and the failure modes.
## Before you start
You need:
- A GoodWe **ES / BP family** inverter (AA55 / RS485 meter-bus generation)
- The vendor's meter-emulating controller **disconnected** from the bus
- A T-CAN485 (ESP32) flashed with `firmware/goodwe-master.yaml`
- A grid-power sensor already working in Home Assistant, updating every ~510 s
## The safety model, in short
The inverter **holds its last command forever** — it has no meter-timeout. So:
- The ESP32 commands 0 W if this add-on stops refreshing for ~30 s, and keeps
commanding it.
- This add-on commands 0 W when its inputs go missing, when you stop control,
and when it shuts down.
- The **optional RS485 e-stop** is the only thing that covers this machine
dying. Without it, a failed host leaves the battery latched at its last
command until someone intervenes.
**If anything looks wrong: stop the add-on.** That commands 0 W and the
hardware holds it there.
## Configuration
### Sources
| option | required | meaning |
|---|---|---|
| `meter_entity` | yes | Net grid power. **Positive must mean importing** |
| `meter_invert` | | Flip the sign if the meter reports the other way |
| `soc_entity` | yes | Battery state of charge — use the **ESP32's own read** |
| `batt_entity` | yes | Battery power — again the ESP32's read, `+` = discharging |
| `batt_invert` | | Flip if needed |
| `setpoint_entity` | yes | The ESPHome `number.*_goodwe_setpoint_w` |
Use the ESP32's readings rather than the inverter's cloud or dongle sensors:
those serve cached values, and a stale reading here ends the maintenance charge
phase having charged nothing.
### P1 meter ingestion
`meter_entity` above expects one signed sensor, which usually means a template
someone wrote by hand. Setting `meter_source` moves the whole derivation into
the add-on, where it is done once and tested, and replaces `meter_entity`
entirely.
Which mode you want depends on what your P1 reader publishes, and there are two
shapes in the wild:
- **Two unsigned registers**, consumption and injection, which is what a Belgian
P1 read over DSMR gives you → `ha_dsmr`, or `mqtt_p1` for a bridge. The add-on
subtracts them.
- **One signed figure**, positive = import and negative = export, which is what
a HomeWizard P1 gives you (`sensor.p1_meter_active_power`) → `ha_signed`. The
add-on splits it. `ha_dsmr` **cannot** read this: it wants two registers and
rejects a negative one outright, which is every exporting telegram.
Either way, do not build the missing shape out of template sensors. The point of
`meter_source` is that the sign convention is derived in one tested place rather
than in YAML nobody reviews underneath a safety input.
| option | default | meaning |
|---|---|---|
| `meter_source` | `off` | `off` keeps `meter_entity`. `ha_dsmr` subscribes to the DSMR integration over the HA WebSocket; `mqtt_p1` reads a topic; `ha_signed` subscribes to one signed entity over the HA WebSocket |
| `meter_phases` | 1 | 1 or 3. Must match the telegram, or every telegram is rejected and logged |
| `meter_max_age_s` | 30 | Beyond this the reading is stale and grid power reads as *missing*. On its own it does **not** command 0 W — see the timing note below. It is also the longest a reading is held forward into the 15-minute average |
| `meter_mqtt_topic` | | `mqtt_p1` only |
| `p1_import_entity` | | The **unsigned** consumption sensor. Do not point this at a signed template |
| `p1_export_entity` | | The **unsigned** injection sensor |
| `p1_phase_import_entities` | `[]` | L1..L3, in order. Needed for the capacity-tariff peak on a three-phase connection |
| `p1_phase_export_entities` | `[]` | L1..L3, in order |
| `p1_net_entity` | | `ha_signed` only. The **signed** net-power sensor: `+` import, `-` export |
| `p1_phase_net_entities` | `[]` | `ha_signed` only. L1..L3, in order, each signed the same way. Needed for the capacity-tariff peak on a three-phase connection |
#### How long a dead meter takes to reach 0 W
`meter_max_age_s` and `stale_input_s` **stack**. They are two different clocks
and neither one is the whole answer:
| step | option | default |
|---|---|---|
| telegrams stop, P1 sample goes stale, grid power starts reading *missing* | `meter_max_age_s` | 30 s |
| inputs have been missing long enough for the loop to command 0 W | `stale_input_s` | 15 s |
| **total, meter death → 0 W commanded by this add-on** | | **45 s** |
So in P1 mode `stale_input_s` is *not* "how long inputs may be missing before
commanding 0 W" measured from the meter dying — it is measured from the moment
the P1 sample already went stale. Size the pair together: the ESP32's own
watchdog commands 0 W after ~30 s of silence from this add-on regardless, and
that layer is unaffected by either option.
There is **no fallback to an inverter-side power figure**, deliberately. The
inverter's own AC power tracks its battery almost perfectly and the real meter
hardly at all, so a controller that failed over to it would be regulating
against its own output while looking healthy.
The `mqtt_p1` payload is one JSON object per telegram, and the schema is strict —
a key it does not recognise is a telegram from something other than what was
tested, and guessing a key here means guessing a kilowatt:
```json
{"import_w": 1234.0,
"export_w": 0.0,
"phases": [{"import_w": 500, "export_w": 0},
{"import_w": 400, "export_w": 0},
{"import_w": 334, "export_w": 0}],
"timestamp": "2026-08-24T18:00:05+02:00"}
```
`phases` and `timestamp` are optional; `timestamp` must carry a UTC offset. Where
it is present it is used for the age, which is what stops a retained message
replayed on reconnect from presenting a ten-minute-old reading as current.
#### `sensor.p1_sample_age_s`
Published over MQTT discovery whenever a broker is available: **seconds since the
newest accepted telegram**, refreshed every second rather than only when a
telegram lands. The ESP32's stale-input watchdog subscribes to this exact entity
id, so do not rename it.
The reason it is recomputed against the clock is that Home Assistant only pushes
a state when the state *changes*. A meter sitting at a genuinely constant reading
emits nothing, which is indistinguishable — to anything watching the value — from
a meter that has died. Watching the age instead separates the two: it climbs when
telegrams stop and resets when they arrive, whatever the reading says.
The entity is only created when `meter_source` is not `off`. With P1 ingestion
disabled there is nothing feeding it, and an age sensor climbing with no ingester
behind it would trip the firmware watchdog on a system that is working fine.
> **Known limit, `mqtt_p1` only.** The age measures *arrival*, not change. On the
> two HA WebSocket paths (`ha_dsmr`, `ha_signed`) that is exactly right: a frozen
> meter emits no `state_changed`, so nothing arrives and the age climbs. On the
> MQTT path a bridge that is stuck republishing its last telegram keeps arriving,
> so the age stays near zero and a frozen meter still looks fresh. Detecting
> *that* needs a change-detector rather than an arrival-detector, and it is not
> in this version. Prefer an HA WebSocket source where both are available.
### Control
| option | default | meaning |
|---|---|---|
| `max_w` | 2000 | Hard limit on what may be commanded. Start low, raise after commissioning |
| `gain` | 0.6 | Correction per cycle. **At the limit — do not raise** |
| `slew_w` | 1000 | Maximum change per cycle |
| `deadband_w` | 15 | Ignore errors smaller than this |
| `target_grid_w` | -10 | What the meter should rest at. Negative = a slight export |
| `step_w` | 10 | Quantisation |
| `saturation_w` | 500 | Divergence that counts as "the inverter is at a limit" |
| `saturation_cycles` | 3 | How many consecutive cycles before freezing. A cycle is one *changed* meter reading, not a fixed period - see the note below. **Do not set to 1** |
| `integrator_max_w` | 0 | Bound on the loop's accumulator, and 0 means "same as `max_w`". Caps how much stale error can be waiting to unwind when the sign flips. **Do not raise it above `max_w`** - the output clamp already bounds what is commanded, so the only thing extra headroom buys is more cycles of wrong-direction power after every saturation event. Lowering it below `max_w` is the useful direction |
| `heartbeat_s` | 10 | Refresh interval; must stay well under the firmware watchdog |
| `stale_input_s` | 15 | How long inputs may be missing before commanding 0 W. In P1 mode this clock starts only *after* `meter_max_age_s` has already expired — the two stack, see "How long a dead meter takes to reach 0 W" |
| `auto_start` | false | Start controlling on boot (only after commissioning) |
#### Saturation is counted in cycles, not seconds
The specification states the saturation window as **"> 10 s"**. This add-on counts
**cycles** instead, and that is a deliberate, accepted deviation rather than an
oversight - the acceptance criterion is not met as literally written.
A cycle here is one *changed* meter reading: the controller only runs the loop when the
meter value differs from the previous poll. At the reference P1's ~5 s update rate the
default of 3 cycles is usually around 15 s, but there is **no guaranteed wall-clock
window** - a meter that repeats the same value stalls the counter for as long as it
repeats.
Two reasons that is acceptable:
- the control law is a pure function with no clock, which is what makes it testable
without hardware, and a seconds-based window would have to live in the controller;
- a stalled counter is a detection-latency limit and not a runaway risk. The condition
that stalls it - an unchanging meter - stops the whole loop, so nothing accumulates
while it is stalled.
If a guaranteed window matters on your site, raise `saturation_cycles` for a fast meter,
and treat the figure as "N meter updates" rather than "N seconds".
#### Why `target_grid_w` is not zero
The deadband is a one-way ratchet: any resting point inside it holds until
something disturbs it. Import and export are **separate registers on the
meter**, so a rest point of +14 W is billed for every second it holds and no
amount of export cancels it - 14 W all day is 0.34 kWh.
Biasing the target below zero moves that residue into the export register,
which is not billed. The resting band becomes `target ± deadband`, so:
| `target_grid_w` | resting band | worst billed leak | export given away |
|---|---|---|---|
| 0 | -15 … +15 W | ~15 W (0.35 kWh/day) | none |
| **-10** | -25 … +5 W | ~5 W (0.12 kWh/day) | ~10 W |
| -15 | -30 … 0 W | none | ~15 W (0.36 kWh/day) |
Set it to `-deadband_w` if injection is worth nothing to you and you would
rather give the energy away than buy it back. Set it to `0` if you are paid
properly for export, or if you are debugging and want the loop centred.
⚠️ This is a **billing** knob, not a speed knob. If import is arriving in
bursts rather than as a trickle, the cause is tracking lag, and this will not
help - see "Why the tuning is what it is".
### Maintenance
| option | default | meaning |
|---|---|---|
| `maintenance_enabled` | false | Enable the monthly cycle |
| `maintenance_interval_days` | 28 | Minimum gap between cycles |
| `maintenance_start_hour` | 10 | Hour of day a due cycle begins |
| `maintenance_discharge_w` | 2500 | Drain rate (exports the surplus) |
| `maintenance_charge_w` | 2500 | Charge ceiling, capped again by peak headroom |
| `maintenance_soc_floor` | 11 | Drain target — stay just above the inverter's own floor |
| `maintenance_soc_target` | 99 | Charge target |
| `maintenance_hold_min` | 120 | Hold at full so the BMS can balance |
### Tariff (all optional)
| option | meaning |
|---|---|
| `peak_forecast_entity` | Quarter-hour demand forecast, for capacity-tariff markets. Empty = no cap |
| `peak_cap_w` | The site's capacity-tariff target |
| `price_now_entity`, `price_avg_entity` | Dynamic tariff. Empty = never force a paid grid top-up |
On a capacity-tariff site the maintenance charge is capped by the headroom left
under `peak_cap_w`, and if the forecast goes over the cap the charge-only clamp
is dropped so the battery can shave the peak instead. Money outranks the
maintenance schedule.
### Site
| option | meaning |
|---|---|
| `estop_fitted` | Whether the RS485 e-stop is installed. Drives the warning banner |
| `log_level` | `trace`/`debug`/`info`/`warning`/`error` |
## The Web UI
The ingress panel shows live values, why the controller is commanding what it
is, and a **Commissioning** checklist that names any problem in words. It also
carries the three buttons: start/stop control, force a maintenance cycle, and
abort one.
## Status entities
If an MQTT broker is available the add-on publishes setpoint, grid power,
battery power, state of charge, maintenance phase and controller status by MQTT
discovery. This is observability only — the controller works fine without a
broker, and MQTT problems can never affect control.