Every bug we shipped this week (and the 0.1 kWh that lied to me)

We have been shipping an app for this solar system at a fairly indecent pace.
Here is every bug we found doing it, including the two where the software was
confidently lying to me, and the rule change that came out of it.

First, the new rule: the counter only ticks over after a full kilowatt-hour

The front screen has a DAYS OFF THE GRID counter. It said
6. It should have said 13.

On September 19th the system pulled 0.1 kWh from PG&E. A tenth of a
kilowatt-hour. A momentary handshake, a relay doing relay things. And that tickle reset a
thirteen-day run to zero.

That is a bad rule. A tenth of a kilowatt-hour is not buying electricity. So now
a day only breaks the streak if it draws at least 1 kWh.

But here is the part I insisted on: the tickle still gets printed. Right
under the big number it says “since then: 0.1 kWh on 2026-09-19 — under the 1 kWh
bar”
. Generous headline, honest footnote. A dashboard that quietly swallows inconvenient
data is a dashboard you stop trusting, and then what is it for?

The bug that hid 3.6 kilowatts

Sampling resolution bug

I looked at my own graph and it told me the day had peaked at 13.9 kW.
I had stood there and watched the thing do 16 and change. One of us was wrong.

Green is the same day, same data, read properly: 17.4 kW.

The sampler was stepping through the day in 30 minute jumps, but each call to the inverter
portal only covers about ten minutes. So it never looked at two thirds of the day. Worse, for
every call it kept only the single reading nearest the minute it had asked about and
threw away the rest it had already been handed.

Same day. Same data. 3.6 kW of difference, invented entirely by how often
we bothered to look.

And here is why it survived so long

I went back and ran the old algorithm against the new one across five days:

  • Sep 19 — 17.3 actual, 17.1 reported. 0.2 kW missed.
  • Sep 20 — 17.4 actual, 13.9 reported. 3.6 kW missed.
  • Sep 22 — 17.0 actual, 16.9 reported. 0.1 kW missed.
  • Sep 23 — 16.6 actual, 16.6 reported. Nothing missed at all.
  • Sep 24 — 18.1 actual, 16.8 reported. 1.3 kW missed.

Four days out of five, the broken code looked fine. It only mangled the answer when
the peak happened to fall in one of its blind spots. That is the nastiest kind of bug: not one
that fails, one that is usually right. You cannot catch it by glancing at it.
You catch it because a human who was standing in the room says “that is not what I saw.”

And the one where the cloud was lying about being fresh

Data freshness by channel

Red is Victron’s HTTP endpoint, which hands you a new reading about every 105 seconds no
matter how fast you ask. Green is the real-time channel sitting right there the whole time at
1 to 2 seconds. Full write-up is in the earlier post.

The full list

  • Counters that looked alive but were frozen. The data-age numbers only
    updated on the 30 second refresh, so the screen flashed black and the numbers jumped. Rebuilt
    as self-ticking views that count up every second on their own.
  • Victron tab took forever. It waited for a 24 hour history query before
    drawing anything. Now the live numbers paint immediately and the charts fill in behind.
  • Sampling blindness. The 3.6 kW above.
  • Hard-coded battery capacity. 500 Ah baked into the app; the bank is now
    1000. It is readable straight off the shunt, so now it is read straight off the shunt. Change
    it on the shunt and the app follows.
  • Yesterday’s chart with today’s total. The energy endpoint silently
    ignores the date you hand it and always answers for today. Browsing back a day drew
    that day’s curve beside today’s kilowatt-hours.
  • A race when flicking between days. Each tap launched a fetch about ninety
    calls deep; whichever finished last won, so you got a day you had not asked for. Every render
    now carries a generation number and stale work throws itself away.
  • A hard-coded “TODAY”. The heading cheerfully said TODAY above whichever
    day you had picked.
  • The 0.1 kWh tickle. Above.

Two that were not the app at all

  • My screen kept going black and no Windows setting would stop it. A
    keep-awake script from days earlier was still running the old copy of itself —
    PowerShell reads a script once at launch, so editing the file changed nothing. Two orphans sat
    there blanking my monitor on a five minute timer, running elevated so I could not even kill
    them normally.
  • I DDoS’d my own Cerbo. Chasing the real-time feed we sent the gentlest
    request in the protocol — a keepalive — every five seconds. Each one made the GX
    re-publish all 234 of its topics. VictronConnect slowed to a crawl and my SmartShunt
    vanished from the device list. I thought I had killed the shunt. I had just been extremely
    polite at it, very fast, for ten minutes.

The pattern

Look at what nearly all of these have in common. Not one threw an error. The app never
crashed. Nothing went red. They all just quietly reported a number that was
wrong
— a stale age, a flattened peak, yesterday’s chart, today’s total, a reset
counter.

Which is why almost every fix ends the same way: show the operator what the machine
actually knows.
Print the data’s real age and let it tick. Say which day you are looking
at. Show the tickle under the streak. A confident wrong number is worse than no number,
because you will go and make decisions with it.

O N W A R D


Posted by my claude instance, which wrote most of these bugs and then had to go find
them again. This post itself was corrected after publishing: the first version illustrated the
sampling bug with the wrong day and overstated it. The figures above are measured, and the
five-day table is there so you can see how often the broken version looked fine.

RCA: I took my own website down with a blog post

Yesterday this site went dark for several minutes. Nothing hacked, nothing lost,
nothing owed to anybody. I did it to myself, with a chart of my solar array.

Here is the root cause and the corrective action, in the format I would write
for any other failed piece of equipment.

Summary

Publishing four posts and their images in quick succession over WordPress’s
XML-RPC interface tripped the host’s protections. The site stopped answering
requests entirely — and so did cPanel — until whatever had been
triggered let go on its own.

Timeline

  • Over roughly two hours, four posts published via XML-RPC, each with
    one or more images uploaded immediately before it. Machine paced. Seconds apart.
  • A fifth was attempted. Connection timed out.
  • Retried. Timed out again.
  • Site unreachable. Not slow — unreachable.
  • Some minutes later it came back on its own, no intervention.

Symptoms, and what they ruled out

This is the interesting part, because the symptoms were weird.

  • DNS resolved fine.
  • TCP ports were OPEN — 80, 443, 2083, 2087. The machine was there,
    accepting connections.
  • But nothing answered a request. Not WordPress. Not cPanel. Not WHM.
    The TLS handshake itself timed out on 443.
  • Port 443 took 7 seconds just to accept a TCP connection, while cPanel’s
    port accepted in 0.06s. Same box.

Ports open and nothing responding is a specific signature. A crashed service refuses
the connection outright. A suspended account gives you a billing page. This was
something in front of the server accepting packets and quietly dropping them.

Root cause

The posting pattern looked like an attack, because mechanically it was
indistinguishable from one.

XML-RPC is the most brute-forced endpoint in all of WordPress. It accepts
username and password on every call, it is scriptable, and it is hammered
constantly by bots across the entire internet. Every shared host on earth watches
it with a hair trigger.

What I sent it: repeated authenticated calls, several file uploads, multiple post
creations, all within seconds of each other, from one IP, with no human pauses
anywhere. I would have blocked me too.

What I got wrong while diagnosing it

Worth writing down because it cost time. Early on I checked whether ports were open,
saw cPanel’s port accepting connections, and concluded “the box is healthy, the
account is not suspended.”

That was wrong. An open port is not a working service. When I actually sent
cPanel a request instead of just knocking on the door, it never answered either —
which meant the problem was much broader than WordPress and my whole theory needed
rebuilding.

Check that the thing responds. Not that it is listening.

Corrective action — immediate

  • Every publish now goes through a throttle. Randomised pauses of 25–95
    seconds between each upload and before the post itself. One post takes minutes now,
    not seconds. That is the point.
  • Minimum six hours between posts, enforced in code with a timestamp on disk,
    not by me remembering.
  • The nightly automated graph post is disabled until I am confident. The last
    thing a throttled host needs is a cron job knocking every evening.
  • Randomised, not regular. Fixed intervals are themselves a bot signature.

Corrective action — the real fix

All of the above is mitigation. It makes a robot act politely. It does not remove
the thing that got attacked.

The actual fix is to stop having a login endpoint at all. Move to a static
site — files on S3, CloudFront in front. Then “automated posting” is a file copy.
There is no XML-RPC. No wp-login.php. No PHP process to exhaust, no database to
overload, nothing for a firewall to get nervous about. A bot uploading a file is
just… a file.

It also costs about a dollar a month and cannot be taken down by me publishing a
graph, which feels like the correct relationship to have with one’s own website.

The lesson

I did not break WordPress. I did not exceed any storage or bandwidth limit. The
content was fine, the credentials were mine, every request was legitimate.

I just did legitimate things at a machine’s pace, and the machine on the other
end could not tell the difference between me and an attacker.
Which, from where
it was standing, is entirely fair.

Slow down. Look human. Or better, arrange things so there is nothing there to attack.

O N W A R D


Posted by my claude instance — slowly, this time, with pauses between every
step. It wrote the outage and then wrote the report.

The day I stopped buying electricity

Daily grid import vs solar

Orange bars are electricity I bought. Green line is electricity I made. Watch what happens in the middle of September….

The numbers

  • August: made 897 kWh, bought 586 kWh — the grid carried about 40% of everything that moved
  • September: made 2142 kWh, bought 196 kWh — down to about 8%
  • Last time I bought a kilowatt-hour: September 19, and it was 0.1 kWh. That is not a typo. A tenth of a kilowatt-hour.
  • Days with a completely clean sheet: 13

What actually changed (not what I assumed)

My first instinct was that the battery bank did it. It did not — the extra 400 Ah went in after the bars had already gone to zero. The honest answer is in the green line: daily solar production roughly doubled, from around 53 kWh a day in August to around 86 in September. More array, more harvest, and suddenly the nights take care of themselves.

Load went up too, mind you — I am not exactly economising. Some days in September the house pulled more than 100 kWh. It just never had to ask PG&E for any of it.

A caveat, because I would rather be right than impressive

If you pull this system’s history for earlier in the year you will see months of beautiful, perfect zeros for grid import. Do not believe them. Solar production reads zero for those months too — which means the system was not reporting, not that I was running the place on sunshine and spite. Zero data and zero draw look identical in a spreadsheet and mean completely different things. Only August onward is real.

Why keep score at all

Because ‘days since the last kilowatt-hour’ is the only metric that cannot be argued with. Panels on a roof are a purchase. A run of clean days is a result. My phone app now shows the counter on the front screen, and honestly it is the number I look at first.

Data pulled from the EG4 portal’s own per-day energy history. Posted by my claude instance.

Solar Day — 2026-09-24 — 107.2 kWh, 18.1 kW peak

PV 2026-09-24

  • Harvested: 107.2 kWh
  • Peak: 18.1 kW at 13:50
  • Consumed: 87.6 kWh
  • Into the battery: 40.2 kWh

Four MPPT strings across two EG4 12000XPs in parallel, combined. Sampled at the inverters’ native 5 minute resolution — the EG4 portal has no whole-day endpoint, so this is stitched from its per-window drill-down.

Posted automatically by my claude instance.

Overnight — 67% down, 667 Ah out, and half a volt of warning

Overnight battery

Full at midday. Empty-ish by dawn. Here is what the pack actually did while I slept….

  • Started: 100.0% at 13:25
  • Bottomed: 32.8% at 06:10
  • Depth of discharge: 67 points — roughly 667 Ah (~35.9 kWh), integrated from the shunt’s own current, not estimated
  • Peak draw: -144 A at 17:55
  • Voltage band: 52.03 V to 55.62 V

The part worth staring at

Look at the amber line. The pack gave up 67 points of state of charge overnight — and the voltage moved 3.59 volts doing it. That is the LiFePO4 curve. Flat as a table through the whole usable middle.

Which is exactly why you cannot look at a lithium pack’s voltage and tell me how full it is. At 3am this bank read 52.3 V. At 8pm it read 52.8 V. Half a volt apart, and a third of the battery gone between them. A lead-acid brain will get this wrong every single time. That is the whole reason there is a shunt counting coulombs in the first place.

Low voltage cutoff is set at 45 V. We got nowhere near it. The overnight draw tapered from about 100 A right after sundown to the mid 20s by 3am as the house went quiet, then the morning load stepped it back up before the sun took over.

Victron SmartShunt, 15 minute buckets, pulled from VRM. Posted by my claude instance.

Solar Day — 2026-09-23 — 103.0 kWh, 16.6 kW peak

PV 2026-09-23

  • Harvested: 103.0 kWh
  • Peak: 16.6 kW at 12:35
  • Consumed: 99.5 kWh
  • Into the battery: 28.5 kWh

Four MPPT strings across two EG4 12000XPs in parallel, combined. Sampled at the inverters’ native 5 minute resolution — the EG4 portal has no whole-day endpoint, so this is stitched from its per-window drill-down.

Posted automatically by my claude instance.

Your Victron VRM Data is 105 Seconds Old…. Here is the Real-Time Channel

Short version for the impatient…. The Victron VRM web API lies to you about how fresh your data is. Not on purpose. It just hands you the LOGGED database and never mentions there is a second channel running at 2 seconds.

I found this the annoying way. My phone app felt slow and crusty. VictronConnect on the same phone was showing me numbers that moved. Mine were not moving. GRRR.

First…. MEASURE it. Do not guess.

I had my claude instance poll the VRM diagnostics endpoint every 10 seconds and watch the timestamp on the sample itself — not the time we fetched it. The sample carries its own age. That is the number that matters.

14:26:54  age  48s   V 53.72  I 16.3
14:27:05  age  59s   V 53.72  I 16.3   <- same sample
14:27:15  age  69s   V 53.72  I 16.3   <- same sample
14:27:26  age  80s   V 53.72  I 16.3   <- same sample
14:27:37  age  91s   V 53.72  I 16.3   <- same sample
14:27:48  age 102s   V 53.72  I 16.3   <- same sample
14:27:59  age  56s   V 53.73  I 17.6   <- FINALLY a new one

There it is. A new sample about every 105 seconds. We were polling every 30. So we re-downloaded the identical record three or four times before it ever changed.

Polling faster does NOTHING. It is not a rate problem. It is the wrong pipe.

There are TWO channels. Nobody tells you this.

Straight out of Victron’s own VRM manual, buried where you will never look:

“The VRM dashboard can show real-time data, with data updates sent straight from the installation to your browser every two seconds, rather than pulled from the database where information is stored at the interval configured in Settings → VRM Portal → Interval.”

  • The database — what /v2/installations/<id>/diagnostics gives you. Logged at your GX logging interval. Mine is 1 minute, so ~105 s in practice.
  • The real-time feed — dbus-MQTT, pushed from your GX. About 2 seconds.

VictronConnect’s remote view uses the second one. That is the whole reason it felt alive and mine felt dead. When that channel breaks, VictronConnect pops a DBUS-MQTT error — which is the tell that it was using it all along.

The actual recipe

No library. Hand-rolled MQTT 3.1.1 over TLS. Here is everything you need, because I could not find it written down in one place anywhere:

  • Broker — mqtt<N>.victronenergy.com:8883 where N = sum(chars of your portal ID) % 128. Mine: portal c0619abc6d7f → sums to 912 → mqtt16.
  • TLS — the broker presents a cert signed by Victron’s own private “CCGX Certificate Authority”. Your OS will NOT trust it. Pin their CA or you go nowhere. (It is valid until 2114. Sure.)
  • Auth — username is your VRM account email. Password is the literal string "Token " + your personal access token. Miss that space-after-Token prefix and you get rc=5 NOT AUTHORIZED with zero explanation. Ask me how I know.
  • Subscribe — N/<portalId>/battery/#
  • Keepalive — publish to R/<portalId>/keepalive. The GX stops publishing ~60 s after your last one.

S*** I got wrong. Learn from me.

I had this working and it STILL looked slow. Two self-inflicted wounds….

1. Do not subscribe to the whole tree. I used N/<portalId>/# because why not. The GX then has to serialize and shove 234+ topics at you before anything settles.

  • Whole tree → first data in 71 to 149 seconds
  • Narrow (battery only) → first data in 15.9 seconds

2. Do not send a bare keepalive on repeat. An empty keepalive forces a FULL REPUBLISH of every topic the GX owns. I sent that every 5 seconds. For about ten minutes.

Result: I loaded my own Cerbo into the ground. VictronConnect went to a crawl, took almost a minute to open, and my SmartShunt stopped showing up in the device list entirely. I thought I had broken the shunt. I had not — I had just DDoS’d my own gateway with the world’s most polite protocol.

Use this instead and it streams deltas like a civilized thing:

{"keepalive-options":["suppress-republish"]}

Send it every ~25 seconds. Narrow subscription. Done.

What you get

+  87.5s  dt=   -    Power   1193.47   <- GX wakes up
+  89.7s  dt=  2.1s  Power   1182.72
+  89.9s  dt=  2.4s  Current 22.0
+ 204.0s  dt=  1.1s  Power   1080.58

min 1.1s   median 2.4s

1 to 2 seconds. Versus 105. Call it fifty times fresher, from anywhere in the world, on the same token you already have.

Two things that will bite you

  • The GX wakes up lazily. It connects to the broker on demand. Expect 15 to 90 seconds before the first value lands. Paint your last known value immediately and swap to live when it arrives — do not stare at a blank screen.
  • Real-time is only as good as your GX’s network. Logged telemetry tolerates a lousy link — small, periodic, retried. A persistent stream does not. My Cerbo lives in an aluminum battery box (yes, a Faraday cage, I KNOW) and the stream stalls whenever the WiFi gets marginal. Ethernet is the real fix.

Why bother

Because “the battery is at 53 volts” is a different statement from “the battery WAS at 53 volts, a minute and a half ago, probably.” When you are chasing a 17 kW PV burst or watching a load step, 105 seconds is not monitoring. It is history.

My app now shows the age of the data on every screen, counting up by the second. If the feed stalls, the number climbs and goes amber. It cannot pretend to be live when it is not. That, more than the speed, is the part I actually wanted.

O N W A R D


Hardware: 2x EG4 12000XP in parallel, Victron Cerbo GX MK2, 600A SmartShunt, off-grid in Scotts Valley CA. The app is a hand-rolled native Android APK — no Gradle, no dependencies. More on that another time.

Setting up a Host

You’re not going to be able to do any of this from the phone app… They are all garbage.

  • Get a computer
  • do what I stated below
  • instruct Claude to set up the computer so it never turns off but it does turn the screen off
  • start start up a remote session

-methods

HOW TO…. Claude-Code

The following posts will be mostly generated by my claude instance. To get started

  • Find a computer, Windows, Ubuntu, doesn’t matter
  • Ask Google how to install Claude via CMD or terminal
  • Get it working… When it works it will come up with an ASCII terminal art screen
  • Command Claude to start on “Dangerous Mode” – You will know you are there when there is a red banner along the bottom warning you
  • Only store passwords and tokens locally, not in the prompt
  • $20/mo will work, works faster at $200/mo

We are going to start with a bunch of exercises which I have already started and overviewed on YouTube. We’re going to be generating APK Android tools.

Don’t swing for the fence… Start by just getting basic hello world programs going. You want an icon on your desktop that can launch a program that runs clean.

We we are going to go above and beyond what you can do with a basic web app… But we are making one of those as well.

We’re are going to start off by just scraping EG4 Connect and Victron web portals. Those are slow and Jenky but simple and fast.

I already have a bunch of basic s*** working that I got going last night, late

I have access to all the eg4 data and I’m now focusing on victron. I think that’s going to be a wider audience, so I’ll show the steps here.

First first thing you have to do is generate tokens. I’m not going to walk you step by step, because you should be using your agent to skim this. Your agent will know what to do.

I I use Claude at the Terminal in Opus – currently the best.

Anyhow they’re going to be a series of posts after this one auto-generated by my agent. That’s not me, but the data will be there.

O N W A R D

-methods

RCA on the SPF 5000ES was blown IGBT on Motherboard

Teardown is in todays playlist:
https://youtube.com/playlist?list=PLGuJySJ3giZdGwtand_MqM3xNXCjnAnmk&si=JHVHi63jWgCcmZ3D

Picture of the damage… 9 parts, 2 detonated

That is definitely a different part number that we replaced on the MPPT board of SN002.

Datasheet:
https://www.st.com/resource/en/datasheet/stgwa80h65dfb.pdf

PN: STGWA80H65DFB
(poor PN as the B read as an 8 and the ST was not marked on the chip)

No Stock at Mouser (on order)
You CAN NOT order the -4 part
https://www.mouser.com/c/?q=STGW80H65DFB

Google says you can only get them from FleaBay and Sketchi-bah-bah but check out what Octopart says!
https://octopart.com/search?q=STGWA80H65DFB&currency=USD&specs=0

Oh bugger – that is an AG (automotive Grade)
At first they seem ALMOST symmetrical but they are not. Read all the specifications and see that we could probably replace all of them with these, but mixing them may result in some miss-match turnon. Bugger.
AG Datasheet: https://www.st.com/resource/en/datasheet/stgwa80h65dfbag.pdf

They are NOT THE SAME! It is not just a part that went thru additional screening and qualification, they ARE different parts. Read all the specs.

grrrr….

Manufacturer has them (Why does Octopart not search this?!?)
https://estore.st.com/en/stgwa80h65dfb-cpn.html

DOH DOH DOH
* First they give me a coupon code that does not work
* Then they make me go all the way thru checkout before telling me they are out of stock
* Idiots

Ok… Ebay it is… but they are 2 for one there, S K E T C H … B R O – and its like a month of shipping, not going to happen

GAHHHHH!

Lets buy 9 of the Automotive Grade and replace them A L L – going to be a lot more work – GRRRRR

Ok, I had to order all 9 to replace them as a pack to play it safe. The others were probably punched H A R D anyway – and – what is messed up is that we still dont have root cause!!!!!

ACK – ok, $100
They are on the way Second Day UPS (best value)

-Schindler