Staff Software Engineer, AI Investigation & Triage

About Command|Link


Command|Link is a global SaaS Platform providing network, voice services, and IT security solutions, helping corporations consolidate their core infrastructure into a single vendor and layering on a proprietary single pane of glass platform. Command|Link has revolutionized the IT industry by tackling the problems our competitors create. In recognition for our unprecedented innovation and dedication, Command|Link was recognized as the SD-WAN Product of the Year, ITSM Visionary Spotlight, UCaaS Product of the Year, NaaS Product of the Year, Supplier of the Year, and the AT&T Strategic Growth Partner. Command|Link has built the only IT platform for scale that solves ISP vendor sprawl and IT headaches. We make it easy for our customers to get more done, maximize uptime and improve the bottom line.


Learn more about us here!


This is a 100% remote position


About your new role:

This role sets the technical direction for the AI-powered layer that sits on top of our alert engine and dramatically reduces mean time to triage. You'll define how LLM reasoning and graph context combine to group alerts into incidents, infer likely narratives, and brief analysts, and you'll make sure that architecture holds up as it's extended across the teams that feed it: alerting and detection, graph and topology, and the analyst-facing investigation experience.


Our edge is correlation: taking diverse sources such as security tooling, monitoring telemetry, syslog, OpenTelemetry, and L2-L4 network protocols, deriving usable network and system topologies and the dependencies between them, and giving that context to LLMs to reason over for correctness, troubleshooting, remediation, and investigation. You'll be the person other engineers come to when a correlation, reasoning, or workflow-durability problem doesn't have an obvious answer yet.


Key Responsibilities:

  • Facilitate the architecture decisions that shape how alert correlation, LLM reasoning, and graph context combine across the investigation and triage pipeline, decisions that other teams (alerting, graph/topology, analyst UX) build on.
  • Own the design for how incidents get grouped, how confidence scores and priority tiers get assigned, and how the AI infers and presents likely incident narratives to SOC/NOC analysts, holding that design to a bar that measurably drives down mean time to triage.
  • Set the technical standard for the collaborative investigation experience: comments, side-bar conversations, shared troubleshooting, and AI-initiated ticketing as a post-triage action.
  • Define the durable workflow architecture, proto-first contracts, deterministic execution, and versioning discipline that this system and adjacent teams' workflows are built on.
  • Step into other teams or projects when a triage-latency, narrative-accuracy, or workflow-durability problem is stuck, contributing hands-on code and design across the alerting, graph-context, and analyst-workbook components as needed, not just within your own team.
  • Balance near-term delivery of MTTT-reducing capability against the long-term technical foundation (workflow durability, model routing, graph schema) that the rest of the org will build on.
  • Actively mentor engineers working across these teams and represent this architecture in conversations with stakeholders outside engineering.
  • Takes on additional responsibilities and projects as needed to support the success of the team and organization.


What you'll need for success:

Required

  • Recognized as an authority in Go and Python, with a track record of building production systems others rely on.
  • Deep, hands-on expertise in LLM tool calling and model routing, and in applying LLM reasoning to structured, real-world problems rather than prototypes.
  • Mastery of Temporal workflow design: determinism, versioning, and proto-first contracts, with experience making these decisions for systems other teams depend on.
  • Demonstrated ability to build novel solutions where no existing playbook applies, ideally including systems that combine graph-based reasoning (e.g., Memgraph or similar) with LLM-driven decision-making.
  • Experience with multi-channel, real-time systems, and comfort operating across a broader technical footprint: Kubernetes/Helm/Docker, AWS/Azure/GCP, Kafka, OpenSearch, and infrastructure-as-code (Argo, Spacelift).
  • A deep, working understanding of the telemetry and protocols this system reasons over, including OpenTelemetry, syslog, SNMP, NetFlow/sFlow, ICMP, and firewall logs, and the judgment to turn that data into usable network and system topologies.
  • Track record of influencing architecture decisions across multiple teams, not just within a single codebase.

Nice to Have

  • Experience with stream processing (e.g., Flink) or detection/anomaly-detection systems (e.g., OpenSearch alerting).
  • Familiarity with endpoint and cloud inventory tooling (osquery, Steampipe) or telemetry pipelines (Vector).
  • Prior experience representing engineering work externally, at conferences, in customer conversations, or in published writing.

 

Why you'll love life at Command|Link

Join us at CommandLink, where you'll have the opportunity to shape the future of business communication. We value the innovative spirit and seek individuals ready to bring their unique vision and expertise to a team that values bold ideas and strategic thinking. Are you ready to make an impact?


  • Room to grow at a high-growth company
  • An environment that celebrates ideas and innovation
  • Your work will have a tangible impact
  • Flexible time off  
  • Fun events at cool locations
  • Employee referral bonuses to encourage the addition of great new people to the team


At CommandLink, we’re committed to creating a fair, consistent, and efficient hiring experience. As part of our process, we use AI-assisted tools to help review and analyze applications. These tools support our recruiting team by identifying qualifications and experience that align with the requirements of each role.


AI tools are used only to assist in the evaluation process — they do not make final hiring decisions. Every application is reviewed by a member of our recruiting or hiring team before any decisions are made.




Software Engineering

India

Czechia

Poland

Chile

Colombia

Brazil

Teilen auf:

NutzungsbedingungenDatenschutzCookiesPowered by Rippling