1-50 of 113 results (3ms)
2025-01-23 §
13:20 <kamila@cumin1002> END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from parse1002 to wikikube-worker1143 [production]
13:18 <kamila@cumin1002> END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming parse1002 to wikikube-worker1143 - kamila@cumin1002" [production]
13:18 <kamila@cumin1002> START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming parse1002 to wikikube-worker1143 - kamila@cumin1002" [production]
13:14 <kamila@cumin1002> START - Cookbook sre.hosts.rename from parse1002 to wikikube-worker1143 [production]
2024-05-31 §
15:44 <cgoubert@cumin1002> conftool action : set/pooled=yes; selector: name=parse1002.eqiad.wmnet,cluster=kubernetes,service=kubesvc [production]
15:43 <claime> pooling and uncordoning parse1002 - T363086 [production]
14:47 <vriley@cumin1002> END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host parse1002.mgmt.eqiad.wmnet with reboot policy FORCED [production]
14:40 <vriley@cumin1002> START - Cookbook sre.hosts.provision for host parse1002.mgmt.eqiad.wmnet with reboot policy FORCED [production]
14:38 <vriley@cumin1002> END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host parse1002.mgmt.eqiad.wmnet with reboot policy FORCED [production]
14:37 <vriley@cumin1002> START - Cookbook sre.hosts.provision for host parse1002.mgmt.eqiad.wmnet with reboot policy FORCED [production]
2024-05-30 §
21:04 <jclark@cumin1002> END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host parse1002.mgmt.eqiad.wmnet with reboot policy FORCED [production]
21:02 <jclark@cumin1002> START - Cookbook sre.hosts.provision for host parse1002.mgmt.eqiad.wmnet with reboot policy FORCED [production]
20:59 <jclark@cumin1002> END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host parse1002.mgmt.eqiad.wmnet with reboot policy FORCED [production]
20:56 <jclark@cumin1002> START - Cookbook sre.hosts.provision for host parse1002.mgmt.eqiad.wmnet with reboot policy FORCED [production]
20:55 <jclark@cumin1002> END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host parse1002.mgmt.eqiad.wmnet with reboot policy FORCED [production]
20:51 <jclark@cumin1002> START - Cookbook sre.hosts.provision for host parse1002.mgmt.eqiad.wmnet with reboot policy FORCED [production]
20:35 <robh@cumin1002> END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host parse1002.mgmt.eqiad.wmnet with reboot policy FORCED [production]
20:32 <robh@cumin1002> START - Cookbook sre.hosts.provision for host parse1002.mgmt.eqiad.wmnet with reboot policy FORCED [production]
20:02 <cdanis@cumin1002> END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host parse1002.mgmt.eqiad.wmnet with reboot policy FORCED [production]
19:57 <cdanis@cumin1002> START - Cookbook sre.hosts.provision for host parse1002.mgmt.eqiad.wmnet with reboot policy FORCED [production]
2024-05-29 §
18:16 <rzl> evacuate cordoned node parse1002: kubectl -n linkrecommendation delete pod linkrecommendation-internal-load-datasets-28616700-7gsqs; kubectl -n linkrecommendation delete pod linkrecommendation-internal-load-datasets-28616700-xl7t4; kubectl -n toolhub delete pod toolhub-main-crawler-28616760-jrhbb # T363086 [production]
10:35 <akosiaris@cumin1002> conftool action : set/pooled=inactive; selector: name=parse1002.eqiad.wmnet [production]
2024-05-27 §
13:15 <hnowlan@cumin1002> conftool action : set/pooled=no; selector: name=parse1002.eqiad.wmnet [production]
2024-04-26 §
11:53 <claime> Silencing all alerts matching parse1002.* for 4 days - T363086 [production]
11:28 <claime> Deactivating puppet for parse1002 - T363086 [production]
2024-04-23 §
14:47 <jclark@cumin1002> END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts parse1002.eqiad.wmnet [production]
14:35 <jclark@cumin1002> START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts parse1002.eqiad.wmnet [production]
14:35 <jclark@cumin1002> END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['parse1002.eqiad.wmnet'] [production]
14:35 <jclark@cumin1002> START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['parse1002.eqiad.wmnet'] [production]
10:43 <jayme> kubectl cordon parse1002.eqiad.wmnet - T363086 [production]
2024-03-06 §
07:55 <akosiaris@cumin1002> conftool action : set/weight=10:pooled=yes; selector: name=(mw1356.eqiad.wmnet|mw1357.eqiad.wmnet|parse1002.eqiad.wmnet|parse1003.eqiad.wmnet|parse1004.eqiad.wmnet|parse1005.eqiad.wmnet|parse1006.eqiad.wmnet|parse1007.eqiad.wmnet|parse1008.eqiad.wmnet|parse1009.eqiad.wmnet|parse1010.eqiad.wmnet|parse1011.eqiad.wmnet|parse1012.eqiad.wmnet|parse1013.eqiad.wmnet|parse1014.eqiad.wmnet|parse1015.eqiad. [production]
2024-03-05 §
16:28 <akosiaris@cumin1002> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host parse1002.eqiad.wmnet with OS bullseye [production]
16:10 <akosiaris@cumin1002> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on parse1002.eqiad.wmnet with reason: host reimage [production]
16:08 <akosiaris@cumin1002> START - Cookbook sre.hosts.downtime for 2:00:00 on parse1002.eqiad.wmnet with reason: host reimage [production]
15:55 <akosiaris@cumin1002> START - Cookbook sre.hosts.reimage for host parse1002.eqiad.wmnet with OS bullseye [production]
2023-07-31 §
10:36 <cgoubert@cumin1001> conftool action : set/pooled=yes; selector: dc=eqiad,name=parse1002.eqiad.wmnet [production]
10:36 <claime> Repooling parse1002 following CPU replacement - T339340 [production]
10:34 <cgoubert@cumin1001> END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for parse1002.eqiad.wmnet [production]
10:34 <cgoubert@cumin1001> START - Cookbook sre.hosts.remove-downtime for parse1002.eqiad.wmnet [production]
09:41 <cgoubert@cumin1001> conftool action : set/pooled=no; selector: dc=eqiad,name=parse1002.eqiad.wmnet [production]
2023-07-25 §
13:42 <cgoubert@cumin1001> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 15 days, 0:00:00 on parse1002.eqiad.wmnet with reason: T339340 - hw troubleshooting [production]
13:42 <cgoubert@cumin1001> START - Cookbook sre.hosts.downtime for 15 days, 0:00:00 on parse1002.eqiad.wmnet with reason: T339340 - hw troubleshooting [production]
13:20 <godog> powercycle parse1002 - T339340 [production]
11:32 <akosiaris> T340087 wikidiff2 rollout done. 1 host is unreachable and will need to be reimaged or upgraded manually to pick this up, parse1002.eqiad.wmnet [production]
2023-07-20 §
04:33 <oblivian@puppetmaster1001> conftool action : set/pooled=inactive; selector: name=parse1002.* [production]
2023-07-19 §
07:54 <_joe_> ran scap pull, pool on parse1002 after powercycling [production]
07:47 <_joe_> powercycling parse1002, console blank, unreachable to network [production]
07:45 <oblivian@cumin1001> conftool action : set/pooled=inactive; selector: name=parse1002.eqiad.wmnet [production]
2023-06-29 §
12:58 <akosiaris@cumin1001> conftool action : set/pooled=yes; selector: name=parse1002.eqiad.wmnet [production]
12:56 <akosiaris@cumin1001> conftool action : set/pooled=no; selector: name=parse1002.eqiad.wmnet [production]