How to make to_string stop at the right time when processing fractional parts

Viewed 98

I wrote a to_string function for my string library and the main part of it looks like this

template<typename T>
inline string num_base(T num,size_t radix,const string radix_table)noexcept{
    string aret;
    do{//do while, also has a return value when num is 0
        T first_char_index{};
        if constexpr(::std::is_floating_point_v<T>)
            first_char_index= ::std::fmod(num,(T)radix);
        else
            first_char_index= num%radix;
        if constexpr(::std::is_floating_point_v<T>)
            num-=first_char_index;
        num/=(T)radix;
        aret.push_front(radix_table[(size_t)first_char_index]);
    }
    while(num);
    return aret;
}
template<typename T>
inline string num_base_mantissa(T num,size_t radix,const string radix_table)noexcept{
    string aret;
    while(num){
        num*=radix;
        T first_char_index;
        num=::std::modf(num,&first_char_index);
        aret+=radix_table[(size_t)first_char_index];
    }
    return aret;
}

The detailed definitions are here, in case anyone still doesn't understand my vague rhetoric

When I tested this function, it worked fine until it ran into (double)1.1:

It outputs "1.100000000000000088817841970012523233890533447265625"

I looked up the reason for this and it seems to be because the underlying binary representation cannot express 1.1, so it has to be approximated instead, but my to_string outputs this approximation as is

I tested std::to_string again and it outputs 1.1 nicely instead of a long string of stuff

I'd like to know how I can modify my function to be less strict as std's version?

1 Answers

Originally from my friend's idea, I now renamed the original to_string to to_string_rough and built the to_string from to_string_rough and from_string_get!

Under the assumption that decimal is used (for ease of example), to_string will first obtain the processing result of to_string_rough and attempt to process the case of list_length consecutive zeros or list_length nines in the decimal part.

  • When consecutive 0s are detected, remove this straight and the part after it
  • When consecutive 9s are detected, remove as above and round up the previous digit

The result is processed by from_string_get, which does a reverse (string-to-floating-point) process and compares it to the original number; if it is full identical, to_string returns the simplified content; if not, to_string continues to process the string until string was end

For more details, see this commit

1.1
1.1(1.100000000000000088817841970012523233890533447265625)
1.12
1.12(1.12000000000000010658141036401502788066864013671875)
1.1314
1.1314(1.1313999999999999612754209010745398700237274169921875)
1.216543215432
1.216543215432(1.21654321543199994692940890672616660594940185546875)
1.3
1.3(1.3000000000000000444089209850062616169452667236328125)

For the tests so far, it works well


Update: After half a day we switched to a faster lookup method (dichotomy) to determine the appropriate value for the truncation loci, but the basic idea remained the same, still requiring a reverse conversion to ensure that truncation would not affect data recovery

The implementation can be seen here

1530.5468561213215646
1530.5468561213215(1530.54685612132152527919970452785491943359375)
1530.5468561213215
1530.5468561213215(1530.54685612132152527919970452785491943359375)
1.12135
1.12135(1.121350000000000068922645368729718029499053955078125)
1.21264
1.21264(1.212639999999999940172301648999564349651336669921875)
54320.215644444444444444444444444445
54320.21564444444(54320.215644444440840743482112884521484375)
21606.1456448565465463218976546
21606.145644856548(21606.14564485654773307032883167266845703125)
21606.145644856548
21606.145644856548(21606.14564485654773307032883167266845703125)

The new method is much smarter and faster


Another update: I've used epsilon to improve the processing of to_string_rough, something like this & this

template<typename T>
inline string num_base_mantissa(T num,size_t radix,const string radix_table)noexcept{
    string aret;
-   while(num){
+   T      epsilon = ::std::numeric_limits<T>::epsilon();
+   while(num >= epsilon){
        num*=radix;
+       epsilon*=radix;
        T first_char_index;
        num=::std::modf(num,&first_char_index);
        aret+=radix_table[(size_t)first_char_index];
    }
    return aret;
}

So now that to_string_rough is mostly as smart as to_string (and there still is no data loss in the conversion), there are still a few times when need to process it further

1.1
1.1(1.1)
1.2
1.2(1.1999999999999999)
1.4
1.4(1.3999999999999999)
1.123456789
1.123456789(1.123456789)
1.1234567891504554
1.1234567891504554(1.1234567891504554)
1.1234567891504554130543435
1.1234567891504554(1.1234567891504554)

Another update: it seems that using epsilon in to_string_rough causes to_string to fail to handle superfine numbers like 0.3000000000000001, although I don't know why, I've stoped the use of epsilon.


Another update: I've added something called information threshold to prevent infinite loops caused by situations like displaying 0.25 in trinary, which I've listed here in case anyone else takes the wrong turn in future

Related