|
2026-07-27
ยง
|
| 10:10 |
<filippo@cloudcumin1001> |
START - Cookbook wmcs.toolforge.k8s.reboot for all NFS workers (T432325) |
[toolsbeta] |
| 10:04 |
<elukey> |
restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 |
[production] |
| 09:58 |
<jiji@cumin1003> |
END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie |
[production] |
| 09:39 |
<elukey> |
restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away |
[production] |
| 09:39 |
<root@cumin1003> |
START - Cookbook sre.mysql.pool pool db1237: Maintenance |
[production] |
| 09:38 |
<jiji@cumin1003> |
END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage |
[production] |
| 09:33 |
<cwilliams@cumin1003> |
dbctl commit (dc=all): 'Depooling db1237 (T431660)', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json |
[production] |
| 09:33 |
<cwilliams@cumin1003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance |
[production] |
| 09:32 |
<jiji@cumin1003> |
START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage |
[production] |
| 09:32 |
<filippo@cloudcumin1001> |
END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component volume-admission (T432325) |
[toolsbeta] |
| 09:30 |
<root@cumin1003> |
END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance |
[production] |
| 09:29 |
<filippo@cloudcumin1001> |
START - Cookbook wmcs.toolforge.component.deploy for component volume-admission (T432325) |
[toolsbeta] |
| 09:17 |
<jiji@cumin1003> |
END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 |
[production] |
| 09:17 |
<jiji@cumin1003> |
END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 |
[production] |
| 09:04 |
<jiji@cumin1003> |
START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 |
[production] |
| 09:04 |
<jiji@cumin1003> |
END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors |
[production] |
| 09:04 |
<jiji@cumin1003> |
START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors |
[production] |
| 09:04 |
<jiji@cumin1003> |
END (PASS) - Cookbook sre.dns.netbox (exit_code=0) |
[production] |
| 09:04 |
<jiji@cumin1003> |
END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" |
[production] |
| 09:04 |
<jiji@cumin1003> |
START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" |
[production] |
| 08:52 |
<jiji@deploy1003> |
helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply |
[production] |
| 08:52 |
<jiji@deploy1003> |
helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply |
[production] |
| 08:52 |
<jiji@deploy1003> |
helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply |
[production] |
| 08:51 |
<jiji@deploy1003> |
helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply |
[production] |
| 08:50 |
<jiji@cumin1003> |
START - Cookbook sre.dns.netbox |
[production] |
| 08:47 |
<jiji@cumin1003> |
START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 |
[production] |
| 08:46 |
<jiji@cumin1003> |
START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie |
[production] |
| 08:44 |
<marostegui> |
Rename tables on s3 T425066 |
[production] |
| 08:43 |
<jiji@cumin1003> |
END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet |
[production] |
| 08:43 |
<root@cumin1003> |
START - Cookbook sre.mysql.pool pool db1203: Maintenance |
[production] |
| 08:43 |
<jiji@cumin1003> |
START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet |
[production] |
| 08:43 |
<jiji@cumin1003> |
START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet |
[production] |
| 08:37 |
<cwilliams@cumin1003> |
dbctl commit (dc=all): 'Depooling db1203 (T431660)', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json |
[production] |
| 08:36 |
<cwilliams@cumin1003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance |
[production] |
| 08:16 |
<root@cumin1003> |
END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance |
[production] |
| 08:11 |
<dcaro@cloudcumin1001> |
END (PASS) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=0) |
[tools] |
| 07:44 |
<phuedx> |
UTC morning backport window done |
[production] |
| 07:37 |
<phuedx@deploy1003> |
Finished scap sync-world: Backport for [[gerrit:1317814|sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) |
[production] |
| 07:28 |
<root@cumin1003> |
START - Cookbook sre.mysql.pool pool db1179: Maintenance |
[production] |
| 07:26 |
<marostegui> |
Rename tables on s3 T426341 |
[production] |
| 07:25 |
<phuedx@deploy1003> |
phuedx: Continuing with deployment |
[production] |
| 07:22 |
<marostegui> |
Drop tables in akwiki nawiki pihwiki - growthexperiments_* T428885 |
[production] |
| 07:22 |
<cwilliams@cumin1003> |
dbctl commit (dc=all): 'Depooling db1179 (T431660)', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json |
[production] |
| 07:22 |
<cwilliams@cumin1003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance |
[production] |
| 07:20 |
<phuedx@deploy1003> |
phuedx: Backport for [[gerrit:1317814|sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. |
[production] |
| 07:16 |
<ryankemper> |
T430880 [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change 1317128, and repooled both; 25/36 hosts complete |
[production] |
| 07:04 |
<phuedx@deploy1003> |
Started scap sync-world: Backport for [[gerrit:1317814|sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] |
[production] |
| 06:57 |
<ryankemper@cumin2003> |
conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet |
[production] |
| 06:56 |
<ryankemper@cumin2003> |
conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet |
[production] |
| 06:40 |
<marostegui@cumin1003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning |
[production] |