Rendered at 19:13:16 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
jesol 1 days ago [-]
Personally I think security is far ahead here compared to normal observability tools. I decided to work on a side-project to try and add SIEM like functionality to a clickhouse backed otel platform; by the end of it the thing I came to believe the tooling blue teams use should be used in observability generally, not just security. Incident/event management is a very powerful concept, and provides a clean framework to hang all of this information on. Then your solution for mechanically finding the neighborhood in k8s is one way to add observations to an event. Some SIEMs have started having agents recommend stuff to be added to an event, which is a nice middle-ground of having agents help but not completely control the discovery and diagnostic effort (as well as providing a clean feedback loop for training data synthesis).
That's all to say, have you considered that framing, and if so, have any opinions on why more general observability tools haven't gone that direction?
jeansilga 7 hours ago [-]
configmap that was updated recently and incorrectly: how would you know the cm update is the cause? How would you sort that out if you made 5/10 updates?
say a node is low on disk: what about setting up monitoring alerts for those king of things.
In general, alerts are of great help. An alert fired after an update is a big smell about that update causing the issue.
That's all to say, have you considered that framing, and if so, have any opinions on why more general observability tools haven't gone that direction?
say a node is low on disk: what about setting up monitoring alerts for those king of things.
In general, alerts are of great help. An alert fired after an update is a big smell about that update causing the issue.