← INDEX Detection engineering with Wazuh · 2 of 4

Wazuh Detection Engineering, Part 2: Decoders and Rules for Your Own Logs

Write Wazuh decoders for plain-text and JSON application logs, build a parent-child rule tree, add frequency rules for brute force, and map every rule to MITRE ATT&CK.

Copy for your AI agent. A condensed version of this entry written as a prompt, so Claude, Cursor, or Copilot can apply it to your own codebase.

View prompt · Custom Wazuh decoders and rules

Task: Add Wazuh detection for a custom application log.

Decide the log format first:

  • JSON lines → <log_format>json</log_format>, no decoder needed, match with <field name="key">.
  • Plain text → <log_format>syslog</log_format> and a custom decoder.

Decoder skeleton (/var/ossec/etc/decoders/local_decoder.xml):

<decoder name="myapp">
  <prematch>^\d+-\d+-\d+T\d+:\d+:\d+\S* \w+ myapp: </prematch>
</decoder>
<decoder name="myapp-auth">
  <parent>myapp</parent>
  <prematch offset="after_parent">auth_failed </prematch>
  <regex offset="after_prematch">user=(\S+) ip=(\S+)</regex>
  <order>user, srcip</order>
</decoder>

Rule skeleton (/var/ossec/etc/rules/local_rules.xml, ids 100000+):

<group name="myapp,">
  <rule id="100100" level="0"><decoded_as>myapp</decoded_as><description>myapp event</description></rule>
  <rule id="100101" level="5"><if_sid>100100</if_sid><match>auth_failed</match>
    <description>myapp: failed login for $(user) from $(srcip)</description>
    <mitre><id>T1110</id></mitre></rule>
  <rule id="100102" level="10" frequency="8" timeframe="120" ignore="300">
    <if_matched_sid>100101</if_matched_sid><same_field>srcip</same_field>
    <description>myapp: brute force from $(srcip)</description>
    <mitre><id>T1110.001</id></mitre></rule>
</group>

Test and deploy:

sudo /var/ossec/bin/wazuh-logtest      # paste a line, check phases 2 and 3
sudo /var/ossec/bin/wazuh-control restart

Rules of thumb: parent rules level 0; one rule per question; regex syntax is Wazuh’s own (\S, \d, \w, \.), not PCRE; same_field needs the field name from the decoder’s <order> or the JSON key.

Part 1 traced an event through the pipeline. Now we make the pipeline understand a log it has never seen: the authentication log of a FastAPI service. By the end there is a rule tree that turns a single failed login into a level 5 alert, eight failed logins from one address into a level 10 brute-force alert with a MITRE technique attached, and a successful login right after a burst of failures into a level 12 alert that somebody should look at now.

The same approach works for any application log. The FastAPI service is just the example that matches the rest of this site.

Two kinds of log, two amounts of work

If the application writes JSON lines, you are most of the way there. The agent ships the line with log_format json, the manager’s built-in JSON decoder flattens every key into a field, and rules match fields directly. No decoder to write.

If the application writes plain text, you need a decoder. A decoder is a small XML block that says “lines that look like this belong to me” and then pulls out named fields with a regex. We will do both, starting with plain text because it teaches how decoders work, then showing how much simpler JSON is.

If you control the application, log JSON

A decoder is a regex you maintain forever. Structured logging from the application removes that maintenance entirely. In Python, structlog or a JSON formatter on the standard library logger takes ten lines. Switch the application before writing decoders, whenever you can.

A plain-text decoder

Suppose the service logs like this:

2026-08-30T09:12:44Z WARN api-gateway: auth_failed user=alice ip=203.0.113.9 path=/login
2026-08-30T09:12:51Z INFO api-gateway: auth_ok user=alice ip=203.0.113.9 path=/login

Decoders live in /var/ossec/etc/decoders/local_decoder.xml. They come in two layers: a parent that claims the line cheaply, and children that do the expensive field extraction only on lines the parent already claimed.

<!-- /var/ossec/etc/decoders/local_decoder.xml -->

<!-- Parent: claim anything from api-gateway. Cheap prematch, no fields. -->
<decoder name="api-gateway">
  <prematch>^\d+-\d+-\d+T\d+:\d+:\d+Z \w+ api-gateway: </prematch>
</decoder>

<!-- Child: failed logins -->
<decoder name="api-gateway-auth-failed">
  <parent>api-gateway</parent>
  <prematch offset="after_parent">auth_failed </prematch>
  <regex offset="after_prematch">user=(\S+) ip=(\S+) path=(\S+)</regex>
  <order>user, srcip, url</order>
</decoder>

<!-- Child: successful logins -->
<decoder name="api-gateway-auth-ok">
  <parent>api-gateway</parent>
  <prematch offset="after_parent">auth_ok </prematch>
  <regex offset="after_prematch">user=(\S+) ip=(\S+) path=(\S+)</regex>
  <order>user, srcip, url</order>
</decoder>

Three details decide whether this works.

offset chains the matching. after_parent means “start looking where the parent’s prematch ended”, and after_prematch means “start the regex where this decoder’s prematch ended”. Without offsets the child regex runs against the whole line and, on a busy log, matches the wrong thing.

The regex dialect is Wazuh’s own, not PCRE. \d digits, \w alphanumerics, \S non-space, \s space, \. any character, \p punctuation, ^ and $ anchors, and (...) captures. There is no \d{4}, no [a-z], no (?:...). If you need PCRE, Wazuh 4.x also accepts <regex type="pcre2">, and the <match> element in rules accepts the same attribute. Use it when the native syntax gets awkward, but expect a small performance cost.

<order> names the captures. The list maps captures to field names in sequence. user, srcip, dstip, url, status, id, action, protocol, data, and extra_data are the conventional names the stock ruleset understands, and using them makes your fields line up with everything else in the dashboard. Custom names also work; <order>tenant, request_id</order> gives you tenant and request_id as dynamic fields.

Restart the manager, open wazuh-logtest, paste the first line, and Phase 2 should print user: 'alice', srcip: '203.0.113.9', url: '/login' under name: 'api-gateway-auth-failed'. If it prints the parent name with no fields, the child’s prematch did not match; check the exact spacing.

The JSON version

Now the same service, logging JSON:

{"ts":"2026-08-30T09:12:44Z","service":"api-gateway","event":"auth_failed","user":"alice","src_ip":"203.0.113.9","path":"/login"}

With <log_format>json</log_format> on the agent’s localfile, no decoder is required. Fields arrive as service, event, user, src_ip, path. The rules below use these names. If you kept plain text, substitute srcip for src_ip and use <decoded_as>api-gateway</decoded_as> in the parent rule instead of the <field name="service"> check.

The rule tree

Rules live in /var/ossec/etc/rules/local_rules.xml. Custom ids start at 100000. Think of the file as a tree, not a list.

<!-- /var/ossec/etc/rules/local_rules.xml -->
<group name="api-gateway,authentication,">

  <!-- 100100: root. Claims every api-gateway event. Level 0 = never alerts on its own. -->
  <rule id="100100" level="0">
    <decoded_as>json</decoded_as>
    <field name="service">api-gateway</field>
    <description>api-gateway event</description>
  </rule>

  <!-- 100101: one failed login -->
  <rule id="100101" level="5">
    <if_sid>100100</if_sid>
    <field name="event">auth_failed</field>
    <description>api-gateway: failed login for $(user) from $(src_ip)</description>
    <mitre>
      <id>T1110</id>
    </mitre>
    <group>authentication_failed,</group>
  </rule>

  <!-- 100102: successful login. Low level, but we need it as a building block. -->
  <rule id="100102" level="3">
    <if_sid>100100</if_sid>
    <field name="event">auth_ok</field>
    <description>api-gateway: login for $(user) from $(src_ip)</description>
    <group>authentication_success,</group>
  </rule>

  <!-- 100110: eight failures from the same address inside two minutes -->
  <rule id="100110" level="10" frequency="8" timeframe="120" ignore="300">
    <if_matched_sid>100101</if_matched_sid>
    <same_field>src_ip</same_field>
    <description>api-gateway: possible brute force from $(src_ip) (8+ failures in 2 min)</description>
    <mitre>
      <id>T1110.001</id>
    </mitre>
    <group>authentication_failures,</group>
  </rule>

  <!-- 100111: a success from an address that just brute-forced. This is the one to page on. -->
  <rule id="100111" level="12" timeframe="600">
    <if_sid>100102</if_sid>
    <if_matched_sid>100110</if_matched_sid>
    <same_field>src_ip</same_field>
    <description>api-gateway: login for $(user) from $(src_ip) after brute force. Possible compromise.</description>
    <mitre>
      <id>T1110.001</id>
      <id>T1078</id>
    </mitre>
  </rule>

</group>

Walk through what each element does.

if_sid makes a rule a child: it is only evaluated if the named rule matched the same event. Rule 100101 is only checked for events that 100100 already claimed, which keeps the whole tree cheap. analysisd evaluates hundreds of thousands of events per second on a modest box precisely because most rules are never reached.

frequency and timeframe turn a rule into a counter. Rule 100110 fires when rule 100101 has matched eight times inside 120 seconds. if_matched_sid is the counting version of if_sid: it looks at previous matches, not the current event.

same_field groups the count by a field. Without it, eight failures from eight different addresses would trip the rule. With same_field src_ip, only eight from one address does. The older same_source_ip element does the same for the conventional srcip field.

ignore suppresses repeats of the same rule for that many seconds after it fires. Without ignore="300", an attacker sending 800 attempts would produce a hundred brute-force alerts. With it, you get one every five minutes while the attack continues, which is what a human wants.

$(field) in descriptions interpolates decoded fields so the alert reads like a sentence. “Failed login for alice from 203.0.113.9” is something an on-call engineer acts on. “Rule 100101 matched” is not.

<mitre><id> attaches ATT&CK technique ids. The dashboard has a whole MITRE view that only works if rules declare techniques. T1110 is brute force, T1110.001 is password guessing, T1078 is valid accounts. Look them up on the ATT&CK site and be specific; a sub-technique is more useful than its parent.1

Level 0 parents are not optional

If you give rule 100100 a level of 3, every single api-gateway event becomes an alert, including the ones your children later refine. Roots are level 0. Only leaves alert. Part 4 relies on this shape to tune volume.

Testing the tree

Restart the manager and feed wazuh-logtest a sequence, not one line. The tool keeps state between lines within a session, so frequency rules can be exercised:

sudo /var/ossec/bin/wazuh-logtest

Paste the failed-login JSON eight times. The first seven should print id: '100101', level: '5'. The eighth prints id: '100110', level: '10', with the brute-force description. Then paste the auth_ok line for the same address and expect id: '100111', level: '12'.

For the plain-text variant, generate the lines with a loop on the agent host and watch alerts arrive:

for i in $(seq 1 8); do
  echo "$(date -u +%FT%TZ) WARN api-gateway: auth_failed user=alice ip=203.0.113.9 path=/login" \
    | sudo tee -a /var/log/api/app.log > /dev/null
done
# on the manager
sudo tail -n 20 /var/ossec/logs/alerts/alerts.json | jq -c 'select(.rule.id | startswith("1001")) | {id: .rule.id, level: .rule.level, desc: .rule.description}'

Patterns you will reuse

A few shapes cover most application detections. Keep them as templates.

Field allowlist. Alert when a field is anything other than an expected value. Wazuh has no “not equals”, so express it as a parent that matches all values and a child at level 0 that catches the allowed ones, leaving the parent to alert on the rest. Alternatively, use <field name="method" negate="yes">^GET$|^POST$</field>, which 4.x supports on field, match, and regex.

Impossible combination. A field that should never appear with another: an admin action from a non-admin role, a payment from an unverified account. One child rule with two <field> elements; all conditions in a rule are ANDed.

Silence after first sight. A rule that should alert once, then stay quiet: use ignore with a long window, or, for permanent suppression per key, <if_matched_sid> with <different_field> to alert only when a new value appears.

Time-of-day. <time>22:00-06:00</time> restricts a rule to a window; <weekday>weekends</weekday> to days. An admin login at 3 a.m. on a Sunday deserves a different level from one on Tuesday afternoon.

Keep rules with the application

The decoder and rules you just wrote belong in the application’s repository, next to the code that produces the log format. When the log format changes, the rules change in the same pull request. Deploy them to the manager with configuration management or a small script that copies the two files and restarts wazuh-manager, and treat the restart as a deployment step with a rollback: keep the previous file and run wazuh-logtest against a fixture of known lines before and after.

Part 3 attaches an active response to rule 100110 that blocks the attacking address, and builds it so that it cannot block the office, the load balancer, or the on-call engineer.

// REFERENCES
  1. Custom decoders Official guide to parent and child decoders, prematch, offsets, and order
  2. Custom rules Rule tree construction, if_sid, if_matched_sid, frequency, timeframe
  3. Rule syntax reference Every rule element and attribute, including same_field, ignore, time, weekday, negate, and mitre
  4. Regex syntax reference The native OS_Regex dialect and when PCRE2 is available
  5. MITRE ATT&CK: T1110 Brute Force The technique and sub-techniques used in this part's rules
  6. Wazuh stock ruleset on GitHub Read the official rules for sshd and web servers as worked examples of the same patterns
  7. structlog Structured logging for Python, the easiest way to avoid writing decoders at all

Footnotes

  1. ATT&CK ids in rules also drive the “MITRE ATT&CK” module in the dashboard, which counts techniques seen per agent. It is one of the few dashboard views that stays useful at scale, but only for rules that declare ids. ↩