it support

Designs and implements technical support processes and systems for information technology environments, including ticketing workflows, knowledge bases, escalation procedures, and user-training materials. Installs, configures, and troubleshoots hardware, software, and network components and analyzes incidents, logs, and service metrics to resolve root causes and improve SLAs and operational reliability.

itsupport

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$170K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Agentic Troubleshooting Guide Automation for Incident Management

Oct 11, 2025
JM
JIAYI MAO
🏛️ Tsinghua University | Microsoft | Microsoft Research

Manual execution of Troubleshooting Guides (TSGs) in large-scale IT systems is inefficient and error-prone, while existing LLM-based approaches struggle with poor TSG quality, complex control flow, data-intensive queries, and parallel execution requirements. Method: We propose an end-to-end automation framework comprising: (i) TSG Mentor to enhance guide quality; (ii) an offline phase leveraging LLMs to construct a structured execution DAG and generate domain-specific Query Preparation Plugins (QPPs); and (iii) an online phase employing a DAG-guided, memory-augmented scheduler and executor that ensures correctness and enables task-level parallelism. Results: Evaluated on real-world TSGs and incidents, our framework achieves a 94% success rate with GPT-4.1—significantly outperforming baselines—and reduces execution time for parallelizable TSGs by 32.9%–70.4%, while also improving token efficiency and latency.

Addressing LLM limitations in handling complex control flow and data queriesAutomating troubleshooting guides to reduce manual execution errors and delaysImproving parallel execution efficiency for IT incident management workflows

为解决ERP系统中数据集成和流程监控的碎片化问题,本文提出一种企业流程控制塔,通过集成状态观测、语义翻译、机器学习诊断等方法提升IT团队的工作效率。

Electronic Data Interchange (EDI)Enterprise Resource Planning (ERP)Intermediate Document (IDoc)

This work addresses the gap in current systems education, where learning resources often consist of superficial tutorials or AI-generated summaries that inadequately convey foundational design principles and thus fail to cultivate robust engineering capabilities. To remedy this, we propose a structured learning pathway centered on seminal research papers from distributed systems, operating systems, and big data domains. Integrating insights from leading academic curricula and industry practices, our approach emphasizes technical depth and problem-solving reasoning. By engaging learners in close reading of original literature, critical analysis of architectural trade-offs, and cross-domain synthesis, the framework fosters a deep understanding of underlying mechanisms and cultivates systems thinking—thereby equipping practitioners to effectively tackle complex engineering challenges and progress toward professional-level systems expertise.

big datadistributed systemsoperating systems

MIT Lincoln Laboratory: A Case Study on Improving Software Support for Research Projects

Dec 01, 2025
DS
Daniel Strassler
🏛️ MIT Lincoln Laboratory

MIT Lincoln Laboratory faced systemic challenges in research-oriented software development—including low engineering efficiency, weak software engineering practices, and poor cross-team collaboration. Method: This study proposes an integrated “tools–talent–governance” framework: (1) a centralized, scalable toolchain; (2) a unified, dynamic talent capability map; and (3) a cross-functional Software Stakeholder Committee. Using organizational behavior analysis, software engineering maturity assessment, platform architecture design, and resource allocation modeling, critical bottlenecks were identified and interventions empirically validated. Contribution/Results: The work yields a reusable, scalable pathway for research software engineering (RSE) adoption. Deployed across multiple high-priority projects, it reduced average software delivery cycle time by 32% and improved requirements alignment accuracy by 47%, accelerating the laboratory’s transition from ad hoc coding to sustainable, engineering-driven scientific software development.

Identifying challenges in efficient research software development processesImproving software engineering culture and effectiveness in research projectsProposing centralized support and staffing solutions for software development

To address the challenges of standardizing Site Reliability Engineering (SRE) practices in heterogeneous environments and balancing system reliability with development agility, this paper proposes a customizable SRE process framework. The framework integrates automated operations, multidimensional observability (metrics, logs, traces), error-budget-driven governance, standardized incident response, and progressive delivery (canary and blue-green deployments). It is designed for cross-technology-stack adaptability, enabling contextual implementation of core SRE principles. Evaluated in production systems, the framework reduced mean time to recovery by 42%, decreased unplanned outages by 67%, lowered operational staffing requirements by 35%, and achieved 99.99% service availability. Its primary contribution is the first methodology for customizing SRE processes specifically for heterogeneous environments, empirically demonstrating synergistic improvements in both system reliability and operational efficiency.

Analyzes SRE processes to boost efficiency, reduce downtimeExplores SRE for scalable, reliable software systemsPresents adaptable SRE techniques for diverse environments

Latest Papers

What's happening recently
View more

This study addresses occupational burnout among Security Operations Center (SOC) practitioners, often stemming from misalignment between job demands and individual capabilities. Drawing on flow theory, the authors conduct an inductive content analysis of 106 global SOC job postings to systematically map the prevalence of certifications (e.g., CISSP), technical skills (e.g., Python, Splunk), and soft skills—particularly communication skills, mentioned in 50.9% of listings. The research reveals, for the first time, a structured pattern in the skill and certification requirements of SOC roles. These findings provide empirical grounding for achieving challenge–skill balance, refining recruitment practices, and guiding professional development. Furthermore, the study advances the discourse on flow-aligned person–job fit and sets the stage for future investigations into the impact of artificial intelligence on SOC workforce dynamics.

burnoutcybersecurity workforceflow theory

This work proposes an intelligent agent-based diagnostic framework leveraging large language models (LLMs) to overcome the limitations of traditional root cause analysis methods, which rely on hard-coded rules, incur high maintenance costs, and are tightly coupled with infrastructure. By integrating a Model Context Protocol (MCP) and a constrained tool space, the framework enables agents to autonomously invoke tools for service querying, dependency retrieval, and multi-source data analysis, facilitating stepwise reasoning to pinpoint root causes. A structured investigation protocol ensures traceable and reproducible inference while maintaining robustness under incomplete or ambiguous information, effectively decoupling the model from underlying infrastructure. This approach lays the foundation for autonomous fault diagnosis and change impact assessment, paving the way for automated remediation and risk prediction, thereby significantly enhancing operational efficiency and system safety.

datacenter infrastructurefailure propagationincident diagnosis

This study addresses the profound transformations in user roles, workflows, and collaboration patterns within enterprise software platforms driven by artificial intelligence, which existing role frameworks—such as the BTP user type matrix—struggle to accommodate. Through 20 expert interviews and a participatory design workshop involving 24 participants, the research employs qualitative methods to investigate structural shifts in developer roles on the SAP Business Technology Platform. Findings reveal three key trends: automation of operational tasks, expanded human-AI collaboration, and increased reliance on agent-based systems. In response, the study argues for a necessary reconfiguration of role taxonomies and governance mechanisms, offering both theoretical grounding and practical guidance for designing and governing AI-native enterprise software.

Artificial Intelligenceenterprise softwarehuman-AI collaboration

This study addresses the lack of systematic guidance for enterprise software teams in choosing between monolithic and microservices architectures. The work proposes a decision-making framework that integrates technical and organizational factors, evaluating the trade-offs of each architecture across dimensions such as scalability, reliability, deployment efficiency, and organizational complexity. The assessment is grounded in system scale, business requirements, operational maturity, and long-term maintainability. Through architectural pattern analysis, a structured evaluation model, and multiple case studies, the authors develop a practical selection methodology tailored to real-world engineering contexts. This approach offers enterprises clear architectural evolution pathways and actionable guidelines aligned with their developmental stages, thereby significantly enhancing the rationality and sustainability of system design decisions.

MicroservicesMonolithic ArchitectureOrganizational Complexity