Post

Why I Stopped Manually Categorizing Server Logs

A journey from manual incident triage to building an automated LLM-based system using FastAPI and Groq. Why I chose to orchestrate AI instead of building it.

Why I Stopped Manually Categorizing Server Logs

It’s 3 AM. Your phone buzzes for the tenth time. The alerts are piling up: “Connection timed out”, “Latency high”, “Service unavailable”. Is it a massive DDoS attack threatening the entire infrastructure, or did a background job just hang? You stare at the raw logs, eyes blurring, trying to make a decision that could save or sink the company’s uptime.

This is the nightmare of manual incident triage. And I realized I couldn’t keep doing it this way.

I used to think doing “AI” meant building a brain from scratch by designing architectures, training models, and burning GPU hours. I was wrong. It’s not about building the brain, it’s about tuning it. It’s about taking a massive, pre-trained intelligence and guiding it to solve your specific problem.

🚀 The Context

I built LLM Incident Triage, a lightweight service using FastAPI that acts as a first line of defense. It takes messy, human-written incident reports and pipes them through Groq (Llama 3.1) to return structured, actionable data: Summary, Category, Urgency, and Recommended Action.

⚠️ The Real Problem

The problem wasn’t just that manual triage was slow. It was that it was brittle.

A human operator at 3 AM makes mistakes. We miss details. We categorize “Database locked” as “Low Priority” because we’re tired. Even worse, we waste time fighting ambiguity instead of fighting the fire. I needed a system that never sleeps and doesn’t get tired.

⚖️ The Decision & Trade-offs

I had a choice: host a small AI model on my own machine, or use a hosted API like Groq.

I chose the API, and here is why. My local machine simply didn’t have the GPU resources to run a model capable of understanding complex technical context. Hosting a small model would have been “free,” but it would have been dumb. By using Groq, I got access to Llama 3.1, a model trained by a tech giant with endless resources. This gave me massive capabilities for a fraction of the cost of buying hardware.

The Trade-off? Privacy. Sending data to an external API is always a risk. I have to trust that the data isn’t being leaked. However, I wasn’t worried about “vendor lock-in.” The beauty of modern AI engineering is that the prompt is the product. If Groq disappears tomorrow, I can switch to OpenAI or Anthropic with minimal code changes.

The Reality Check: Hallucinations. AI isn’t perfect. It can confidently lie. That’s why I treat this system as a copilot, not an autopilot. A “human in the loop” is essential to verify the AI’s triage before any drastic automated actions are taken. We trust, but verify.

📊 The Honest Outcome

So, does it work? Yes. The system parses text faster than any human could. It turns “internet lambat” into a structured ticket categorized as “Network Issue” with “Medium Urgency” instantly.

But it’s not magic. It hasn’t failed catastrophically yet, but I know the risk is there. LLMs can hallucinate. In a production environment, I can’t just trust it blindly. I anticipate that my future work will be less about coding features and more about prompt engineering. This involves constantly testing and refining instructions to ensure the AI doesn’t categorize a security breach as a minor bug.

💡 The Lesson

This project changed how I see software engineering. I entered this wanting to be an AI Architect, building the engine. I left as an AI Orchestrator, driving the car. We don’t need to have the resources of a research lab to build powerful AI tools, we just need to know how to ask the right questions.

This post is licensed under CC BY 4.0 by the author.