Rule Cascade
LearnValues and types

Strings and Unicode

Strings are sequences of Unicode code points. len and substring count code points, never bytes or UTF-16 units.

A string is a sequence of Unicode code points. len, substring and . in a pattern count code points. So "café😀" has 5, in every language — even in JavaScript and Java, which store the emoji as two UTF-16 units. Comparisons are exact: case and accents count.

Syntax

strings
{ op: len, args: [{ var: data.nickname }] }               # code points
{ op: substring, args: [{ var: data.name }, 0, 3] }       # the first 3 code points

Example

A nickname has at most 5 code points. "Zoë 👋🎉" has 6. The compute rule fills /username with the first three code points of the name.

strings.ruleset.yaml
ruleCascade: 1.0.0
kind: RuleSet
metadata: { id: learn.strings, version: 1.0.0, title: "Strings and Unicode" }
scope:
  - { level: organization, id: learn }
entities:
  Customer:
    schema: { $ref: "./learn.openapi.yaml#/components/schemas/Customer" }
rules:
  - id: customer.username.default
    kind: compute
    target: { entity: Customer, field: /username }
    operations: [create]
    assign:
      - { field: /username, value: { op: substring, args: [{ var: data.name }, 0, 3] } }
  - id: customer.nickname.short
    kind: validation
    target: { entity: Customer, field: /nickname }
    operations: [create]
    when: { op: exists, args: [{ var: data.nickname }] }
    assert: { op: lte, args: [{ op: len, args: [{ var: data.nickname }] }, 5] }
    severity: error
    finding: { code: LRN-STR-001, message: customer.nicknameTooLong, args: { length: { op: len, args: [{ var: data.nickname }] } } }
messages:
  en:
    customer.nicknameTooLong: "A nickname has at most 5 characters; this one has {length}."
tests:
  - name: six code points are too many
    entity: Customer
    operation: create
    given:
      data: { name: "Zoë Ann", nickname: "Zoë 👋🎉" }
    expect:
      decision: deny
      findings:
        - { rule: customer.nickname.short, fields: [/nickname] }
      effects:
        - { type: value, field: /username, value: "Zoë" }
  - name: five code points are fine, even with an emoji
    entity: Customer
    operation: create
    given:
      data: { name: Ana, nickname: "café😀" }
    expect: { decision: allow, findings: [] }
request.json
{
  "entity": "Customer",
  "operation": "create",
  "data": {
    "name": "Zoë Ann",
    "nickname": "Zoë 👋🎉"
  }
}

Result, from the engine

Decisiondeny1 finding, server channel

  • LRN-STR-001errorblockingA nickname has at most 5 characters; this one has 6./nickname
  • computed value /username = Zoë
Try it YourselfOpens this ruleset and request in the playground. Nothing to install.

Common mistakes

  • Checking a length in the browser with string.length. It counts UTF-16 units, so "😀".length is 2. The engine says 1; use the engine.
  • Expecting lower and upper to change É. They change only the ASCII letters A–Z and a–z.
  • Comparing text that a keyboard may write in two ways (é as one code point, or e plus an accent). Normalise input before you evaluate.

Exercise

Lower the limit to 3 code points. Check that "café" is denied with the message "A nickname has at most 3 characters; this one has 4."

Hint

Change the limit in the assert and in the message. Then make the denied test use a 4-code-point nickname.

Show answer
strings.ruleset.yaml
ruleCascade: 1.0.0
kind: RuleSet
metadata: { id: learn.strings, version: 1.0.0, title: "Strings and Unicode" }
scope:
  - { level: organization, id: learn }
entities:
  Customer:
    schema: { $ref: "./learn.openapi.yaml#/components/schemas/Customer" }
rules:
  - id: customer.username.default
    kind: compute
    target: { entity: Customer, field: /username }
    operations: [create]
    assign:
      - { field: /username, value: { op: substring, args: [{ var: data.name }, 0, 3] } }
  - id: customer.nickname.short
    kind: validation
    target: { entity: Customer, field: /nickname }
    operations: [create]
    when: { op: exists, args: [{ var: data.nickname }] }
    assert: { op: lte, args: [{ op: len, args: [{ var: data.nickname }] }, 3] }
    severity: error
    finding: { code: LRN-STR-001, message: customer.nicknameTooLong, args: { length: { op: len, args: [{ var: data.nickname }] } } }
messages:
  en:
    customer.nicknameTooLong: "A nickname has at most 3 characters; this one has {length}."
tests:
  - name: four code points are too many now
    entity: Customer
    operation: create
    given:
      data: { name: Ana, nickname: "café" }
    expect:
      decision: deny
      findings:
        - { rule: customer.nickname.short, fields: [/nickname], message: "A nickname has at most 3 characters; this one has 4." }
  - name: three code points are fine
    entity: Customer
    operation: create
    given:
      data: { name: "Zoë Ann", nickname: "Zoë" }
    expect: { decision: allow, findings: [] }
request.json
{
  "entity": "Customer",
  "operation": "create",
  "data": {
    "name": "Ana",
    "nickname": "café"
  }
}

Result, from the engine

Decisiondeny1 finding, server channel

  • LRN-STR-001errorblockingA nickname has at most 3 characters; this one has 4./nickname
  • computed value /username = Ana
Course overview

On this page