ascii-chat 0.11.33
Video chat in your terminal
Loading...
Searching...
No Matches
utf8.h File Reference

🔤 UTF-8 Encoding and Decoding Utilities More...

Go to the source code of this file.

Functions

int utf8_decode (const uint8_t *s, uint32_t *codepoint)
 Decode a UTF-8 sequence to a Unicode codepoint.
 
int utf8_display_width (const char *str)
 Calculate terminal display width of a UTF-8 string.
 
int utf8_display_width_n (const char *str, size_t max_bytes)
 Calculate terminal display width of a UTF-8 string up to byte limit.
 
bool utf8_is_valid (const char *str)
 Check if a string is valid UTF-8.
 
bool utf8_is_ascii_only (const char *str)
 Check if a string contains only ASCII characters.
 
size_t utf8_char_count (const char *str)
 Count UTF-8 characters (not bytes)
 
size_t utf8_to_codepoints (const char *str, uint32_t *out_codepoints, size_t max_codepoints)
 Convert UTF-8 string to array of Unicode codepoints.
 
int utf8_next_char_bytes (const char *str, size_t max_bytes)
 Get byte length of next UTF-8 character.
 
int utf8_continuation_bytes_needed (unsigned char first_byte)
 Determine how many additional bytes are needed to complete a UTF-8 character.
 
int utf8_read_and_insert_continuation_bytes (char *buffer, size_t *cursor, size_t *len, size_t max_len, int continuation_bytes, int(*read_byte_fn)(void))
 Read continuation bytes and insert them into buffer at cursor position.
 
const char * utf8_strcasestr (const char *haystack, const char *needle)
 Case-insensitive substring search with full Unicode support.
 
void utf8_write_lowercase (FILE *output, const char *text)
 Write UTF-8 string to file with ASCII lowercase conversion.
 

Detailed Description

🔤 UTF-8 Encoding and Decoding Utilities

This header provides simple, efficient UTF-8 validation and decoding without external dependencies. The implementation handles multi-byte UTF-8 sequences and validates encoding correctness.

CORE FEATURES:

  • Multi-byte UTF-8 sequence decoding (1-4 bytes)
  • Unicode codepoint extraction
  • UTF-8 validation during decoding
  • No external dependencies (pure C implementation)
  • Safe handling of invalid sequences

UTF-8 ENCODING:

UTF-8 encodes Unicode codepoints using 1-4 bytes:

  • 1 byte: ASCII characters (0x00-0x7F)
  • 2 bytes: Latin-1 supplement, etc. (0x80-0x7FF)
  • 3 bytes: Most CJK characters (0x800-0xFFFF)
  • 4 bytes: Rare characters, emoji (0x10000-0x10FFFF)

VALIDATION:

The decoder validates UTF-8 sequences during decoding:

  • Checks for valid byte sequences according to UTF-8 rules
  • Detects overlong encodings (security feature)
  • Detects invalid byte patterns
  • Returns error on invalid sequences
Note
This is a minimal implementation for basic UTF-8 handling.
For full Unicode support, consider using a library like ICU.
Invalid sequences return -1 to indicate error.
Author
Zachary Fogg me@zf.nosp@m.o.gg
Date
October 2025

Definition in file utf8.h.