SOFTSWISS is hiring a Monitoring Systems Engineer to join our Team. We are looking for a detail-oriented engineer to design, maintain, and enhance our monitoring and observability ecosystem, ensuring the reliability, performance, and visibility of critical services across our technology landscape.
Purpose of the role:You will be responsible for building and evolving the monitoring and observability platform that enables teams to detect, troubleshoot, and prevent issues across our production environment. By developing reliable monitoring solutions, improving system visibility, and collaborating with engineering teams, you will help ensure the stability, performance, and availability of our services at scale.
Key responsibilities:Offering on-duty service coverage, encompassing day and night shifts.
Addressing incidents by troubleshooting and resolving issues, even seeking assistance from third-party or vendor support when necessary.
Directing issues or queries to the relevant department as needed.
Keeping detailed records and documentation of current infrastructure challenges and Root Cause Analyses (RCAs).
Contribute to safe and effective internal practices for AI usage in monitoring and incident response workflows.
Collaborating with other teams to understand and define their monitoring needs, then implementing the right solutions.
Setting up and adjusting the monitoring/observability systems for various teams.
Designing and tweaking alerts and dashboards to suit specific needs.
Refining alerts to reduce irrelevant notifications and increase their significance.
Enhancing dashboards for better clarity, understanding, and a more comprehensive view.
Building and sustaining connections between the monitoring systems and other platforms like Jira, Opsgenie, etc. when required.
Establishing and updating a Knowledge Base, covering system configurations, alert processes, troubleshooting guidelines, and user manuals.
Staying updated with the newest trends and best practices to continuously uplift our organization's monitoring capabilities.
Identify opportunities to automate repetitive monitoring and support tasks, including with AI-assisted approaches where suitable.
Minimum of 3 years experience as a Systems Engineer, SRE, DevOps, or Monitoring Support Engineer (L2+).
Good understanding of Linux-like operating systems (Debian-based).
Experience with containerization, virtualization, and orchestration (LXC/LXD, Docker, Kubernetes).
Development experience in any scripting language (Bash, Python, Go, etc) and familiarity with REST API.
Knowledge of basic database concepts (experience with PostgreSQL is preferable), including transactions and WAL.
English proficiency at an Intermediate (B1) level or higher. It's crucial to understand technical terminology related to our specific tech stack and to be able to interpret technical documentation.
Russian proficiency at an Upper-Intermediate (B2) level or higher.
Practical interest in using AI-assisted tools for troubleshooting, automation, documentation, and operational efficiency.
Ability to critically evaluate AI-generated output and validate it before using it in production environments.
Understanding of the risks and limitations of AI usage in infrastructure and production operations.
Private health insurance
Sports benefits
Comprehensive Mental Health Program
Free English lessons (online)
Local language courses
Paid time off
Maternity leave support
Referral program rewards
Upskilling, internal workshops, and participation in professional conferences and corporate events
