Skip to content

STATA writer assumes non-unicode characters for byte count #382

Description

@FlipperPA

I ran into an issue where READSTAT was writing 999 rows instead of 1000.

The row being dropped included a field called description which contained this string:

100 milyondan fazla indirilen letgo'da elektronikten arabaya, anne & bebekten ev eşyalarına kadar ihtiyacın olan her şeyi bulabilir, kullanmadığın eşyaları kolayca satabilirsin. "Cüzdanım Güvende" ile aradığın ürünler taksit avantajıyla kapına kadar gelir; ilanların tüm Türkiye'ye ulaşır. Bizce hiçbir hayal ikinci planda kalmamalı. İkinci el alışveriş sana bu hayallerini yaşaman için bir fırsat yaratmalı. Sıfır almak yerine letgo'yla kazanmayı, bilinçli alışverişe ve hayallerimize yatırım yapmayı seçiyoruz. Ne kendi hayatımızı ne de dünyayı daraltıyoruz. Bu yüzden, "letgo'yla kazan, hayallerini yaşa" diyoruz! Eşyalara ikinci bir şans vermeye inanıyoruz! İhtiyacın olanı uygun fiyata bulurken kullanmadığın eşyaları da kolayca kazanca çevirip hem bütçeni rahatlatıyor hem de dünyayı korumana yardımcı oluyoruz. Türkiye'de ikinci el alışverişini baştan aşağı yeniliyoruz. Güçlü, gelişen bir kültürde, etki yaratan işlere odaklanıyoruz. Kullanıcı deneyimine tutkuyla bağlı, yenilikçi ve fark yaratmak isteyen biriysen seni de aramızda görmeyi çok isteriz. ----- letgo is the go-to platform for buying and selling second-hand items, with over 100 million downloads and millions of listings. We simplify local transactions for everything from electronics to cars, baby items to home goods. We believe no dream should take second place. Second-hand shopping with letgo is the smart choice, fueling your aspirations and expanding your world by making space for what truly matters—your dreams. Why buy new when you can save and achieve more? That's why we say: "Live your dreams!" We champion the circular economy, reducing waste by connecting users with what they need. Our "Cüzdanım Güvende" feature offers secure payments and shipping. We're shaping the future of second-hand in Turkey, creating a beloved user experience. We value passionate individuals who thrive on impactful work and new challenges within our evolving culture.

While the character count is 1,946, the UTF-8 byte count is 2,067. So it fits under a 2,000-character assumption, but it does not fit into str2045's bytes. STATA's older fixed string types are str1 through str2045; longer text should use strL. STATA’s own docs also note that UTF-8 characters can take 2, 3, or 4 bytes, so character count and byte count diverge for Unicode-heavy text.

I'll work on a fix that addresses this with proper counting.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions