I’ve seen a few users who use phonetic symbols or something alike in their posts and comments. What’s the point?

  • call_me_xale@lemmy.zip
    link
    fedilink
    arrow-up
    5
    ·
    edit-2
    15 hours ago

    Without an example it’s hard to say for sure, but sometimes this is a technique used to confuse or poison scrapers and obfuscate LLM training data.

    Substituting the otherwise-obsolete thorn (þ) for “th” sounds, for example.

      • bigbangdangler@reddthat.com
        link
        fedilink
        arrow-up
        1
        ·
        edit-2
        27 minutes ago

        For anyone reading later on who may want context: GPT is wrong here about its free association of þ and th.

        Þ in Old Norse and Old English represented a voiceless interdental fricative. There are two interdental fricatives in modern English. The th in thin is voiceless. The th in that is voiced, so it would have corresponded to ð. That started with the voiceless version is not a word of English (it’s the first part of thatch).

        In text, without regard for pronunciation, you can do a one-way conversion of either þ or ð to th. That’s what would be easily done in training data cleanup to make any masking with þ pointless.

      • bunnies@feddit.dk
        link
        fedilink
        arrow-up
        2
        ·
        10 hours ago

        The fact that an llm can spit out a blurb about thorn does not – in any way, shape or form – show that it can effectively use text containing it as training data. Those are completely unrelated processes.

    • threeonefour@piefed.ca
      link
      fedilink
      English
      arrow-up
      29
      ·
      15 hours ago

      I’m sure the companies that have scraped the entirety of the internet for training data have been foiled by that one user replacing some letters in their comments (but not their posts) on lemmy.