The regex ^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$ matches standard UUID strings. It enforces the exact structure of eight hexadecimal digits, followed by four groups of four hexadecimal digits, separated by hyphens. This pattern ensures the string contains only lowercase hexadecimal characters and follows the canonical 8-4-4-4-12 grouping.
Why Standard UUID Regexes Fail
Many developers copy-paste overly complex regexes for UUIDs, often including unnecessary lookaheads or case-insensitive flags that slow down processing. A common mistake is allowing uppercase letters when the system expects lowercase, or failing to anchor the pattern, which lets partial matches slip through. Another frequent error is ignoring the hyphen positions, resulting in patterns that accept malformed strings like 1234567890abcdef1234567890abcdef.
The goal is a strict validator that accepts only the canonical form. This means exactly 32 hexadecimal digits, divided into five sections by hyphens. The pattern must reject strings that are too short, too long, or contain invalid characters like g, z, or spaces. By focusing on the strict structure, you avoid edge cases where a loose regex accepts invalid data that breaks downstream parsers.
The Recommended Pattern
The most reliable regex for validating a standard UUID is concise and explicit. It relies on character classes for hexadecimal digits and fixed-length quantifiers for each segment. This approach is faster than using alternation or complex lookarounds because the engine simply checks character counts and types sequentially.
Here is the exact pattern to use:
^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$
This pattern assumes lowercase hex digits. If your system accepts uppercase, use [0-9a-fA-F] instead. However, most modern systems normalize UUIDs to lowercase upon creation, so sticking to lowercase is often safer for consistency. The anchors ^ and $ ensure the entire string matches, preventing trailing whitespace or extra characters from passing validation.
Step-by-Step Breakdown
Understanding how this pattern processes a string helps when debugging validation issues. Let’s walk through the validation of a specific UUID: 123e4567-e89b-12d3-a456-426614174000.
- **Start Anchor (
^)**: The engine ensures matching begins at the very start of the string. No leading spaces or characters are allowed. - **First Group (
[0-9a-f]{8})**: Matches exactly eight hexadecimal characters. In our example, it matches123e4567. Each character must be a digit (0-9) or a lowercase letter (a-f). If the eighth character is ag, the match fails immediately. - **Hyphen (
-)**: Matches a literal hyphen. This separates the first group from the second. - **Second Group (
[0-9a-f]{4})**: Matches exactly four hexadecimal characters. Here, it matchese89b. - **Hyphen (
-)**: Matches the next literal hyphen. - **Third Group (
[0-9a-f]{4})**: Matches exactly four hexadecimal characters. Here, it matches12d3. - **Hyphen (
-)**: Matches the next literal hyphen. - **Fourth Group (
[0-9a-f]{4})**: Matches exactly four hexadecimal characters. Here, it matchesa456. - **Hyphen (
-)**: Matches the final literal hyphen. - **Fifth Group (
[0-9a-f]{12})**: Matches exactly twelve hexadecimal characters. Here, it matches426614174000. - **End Anchor (
$)**: Ensures no characters remain after the twelfth digit of the last group. Any extra character causes the match to fail.
When you paste this regex and the test string into RegexBuilder, the live highlight shows each segment lighting up in sequence. The capture-group inspector displays the matched segments, confirming that each block meets its specific length requirement. This visual feedback helps verify that the hyphens are correctly positioned and that no segment is too short or too long.
Handling Different UUID Versions
UUIDs come in different versions, primarily Version 1 (time-based) and Version 4 (random). The regex above works for both because it only validates structure, not the specific bits inside. However, some strict systems require checking the version bit. Version 4 UUIDs always have a 4 in the first position of the third group. Version 1 UUIDs have a 1 in the same position.
If you need to enforce Version 4 specifically, modify the third group to require the first character to be 4:
^[0-9a-f]{8}-[0-9a-f]{4}-4[0-9a-f]{3}-[89ab][0-9a-f]{3}-[0-9a-f]{12}$
Note that the fourth group starts with [89ab] because Variant 1 UUIDs require the first bit of the fourth group to be 8, 9, a, or b. This stricter pattern rejects UUIDs that don’t conform to RFC 4122 Version 4 standards. For most applications, the general pattern is sufficient because UUIDs are treated as opaque identifiers. Only add version checks if your backend logic depends on the UUID generation method.
Testing with Live Highlights
Manual testing is tedious and error-prone. Using a tool with live highlighting speeds up verification significantly. When you enter the pattern and test string, you should see distinct colored segments corresponding to the matched parts. If a segment fails, the highlight stops at that point, showing exactly where the mismatch occurs.
For example, if you test 123e4567-e89b-12d3-a456-42661417400 (missing one digit at the end), the highlight will cover everything up to the last group but fail to close the final anchor. The inspector will show that the last group only matched 11 characters instead of the required 12. This immediate feedback helps you adjust the quantifiers or check your input data without guessing.
You can also test edge cases like uppercase letters. If your pattern uses [0-9a-f] and you test 123E4567-E89B-12D3-A456-426614174000, the match will fail at the first uppercase letter. Switching to [0-9a-fA-F] makes it pass. This quick iteration ensures your regex matches your actual data format without over-complicating the logic.