|
2026-09-03
ยง
|
| 09:17 |
<mvernon@cumin2003> |
START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage |
[production] |
| 09:15 |
<topranks> |
put traffic on Lumen codfw<->eqiad link as it is stable T435810 |
[production] |
| 09:14 |
<slyngshede@cumin1003> |
END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams |
[production] |
| 09:09 |
<elukey@cumin1003> |
START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm |
[production] |
| 09:08 |
<elukey@cumin1003> |
END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm |
[production] |
| 09:06 |
<slyngshede@cumin1003> |
START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams |
[production] |
| 09:05 |
<slyngshede@cumin1003> |
END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs |
[production] |
| 09:03 |
<marostegui> |
Move s6 sanitarium from db1165 to db1279 T434775 |
[production] |
| 09:03 |
<mvernon@cumin2003> |
START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm |
[production] |
| 08:58 |
<marostegui@cumin1003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6 |
[production] |
| 08:57 |
<slyngshede@cumin1003> |
START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs |
[production] |
| 08:57 |
<slyngshede@cumin1003> |
END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw |
[production] |
| 08:55 |
<blake@cumin1003> |
START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw |
[production] |
| 08:49 |
<slyngshede@cumin1003> |
START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw |
[production] |
| 08:49 |
<elukey@cumin1003> |
END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage |
[production] |
| 08:45 |
<slyngshede@cumin1003> |
END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru |
[production] |
| 08:45 |
<elukey@cumin1003> |
START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage |
[production] |
| 08:43 |
<elukey@cumin1003> |
END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART |
[production] |
| 08:42 |
<marostegui> |
Move s5 sanitarium from db1161 to db1275 T434776 |
[production] |
| 08:39 |
<marostegui@cumin1003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5 |
[production] |
| 08:38 |
<slyngshede@cumin1003> |
START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru |
[production] |
| 08:37 |
<elukey@cumin1003> |
START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART |
[production] |
| 08:37 |
<jnuche@deploy1003> |
Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): T436812 to prod server (duration: 01m 21s) |
[production] |
| 08:37 |
<elukey@cumin1003> |
END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART |
[production] |
| 08:36 |
<jnuche@deploy1003> |
Started deploy [releng/jenkins-deploy@b146ae7] (releasing): T436812 to prod server |
[production] |
| 08:34 |
<elukey@cumin1003> |
START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART |
[production] |
| 08:33 |
<jnuche@deploy1003> |
Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): T436812 to backup server (duration: 01m 28s) |
[production] |
| 08:32 |
<jnuche@deploy1003> |
Started deploy [releng/jenkins-deploy@b146ae7] (releasing): T436812 to backup server |
[production] |
| 08:27 |
<elukey@cumin1003> |
START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm |
[production] |
| 08:26 |
<elukey@cumin1003> |
END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet |
[production] |
| 08:11 |
<moritzm> |
uploaded wmf-laptop 1.0.7 to apt.wikimedia.org |
[production] |
| 08:03 |
<marostegui> |
Move s2 sanitarium from db1156 to db1271 T434287 |
[production] |
| 07:59 |
<elukey@cumin1003> |
START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet |
[production] |
| 07:59 |
<elukey@cumin1003> |
END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet |
[production] |
| 07:59 |
<elukey@cumin1003> |
END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet |
[production] |
| 07:58 |
<marostegui@cumin1003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2 |
[production] |
| 07:49 |
<elukey@cumin1003> |
START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet |
[production] |
| 07:39 |
<elukey@cumin1003> |
START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet |
[production] |
| 07:29 |
<chlod> |
UTC morning backport window done |
[production] |
| 07:27 |
<chlod@deploy1003> |
Finished scap sync-world: Backport for [[gerrit:1333872|core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s) |
[production] |
| 07:22 |
<chlod@deploy1003> |
chlod, tryvix1509: Continuing with deployment |
[production] |
| 07:22 |
<XioNoX> |
push pfw policies - T436729 |
[production] |
| 07:20 |
<chlod@deploy1003> |
chlod, tryvix1509: Backport for [[gerrit:1333872|core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. |
[production] |
| 07:15 |
<chlod@deploy1003> |
Started scap sync-world: Backport for [[gerrit:1333872|core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] |
[production] |
| 07:15 |
<marostegui> |
Power off db1228 for maintenance |
[production] |
| 07:13 |
<marostegui@cumin1003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance |
[production] |
| 07:01 |
<arnaudb@dns1006> |
END - running authdns-update |
[production] |
| 06:58 |
<arnaudb@dns1006> |
START - running authdns-update |
[production] |
| 06:54 |
<jmm@cumin2003> |
END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public |
[production] |
| 06:52 |
<jmm@cumin2003> |
START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public |
[production] |