Julia implementation for converting string to snake_case/CamelCase

Viewed 296

I am looking for an implementation of the python library https://pypi.org/project/stringcase/ in Julia.

I found the following packages, but they seem all a bit outdated:

Is there an up-to-date Julia library for converting strings to snake_case, CamelCase etc?

Edit: I have the following use-case:

I receive a JSON from a C# framework that uses CamelCase naming convention, which I load into a DataFrame. The resulting DataFrame has column names like: timeStamp, askBestVolume, askBest30MWPrice. I'd like to convert the column names to the snake_case naming convention, i.e.

"timeStamp" => "time_stamp"
"askBestVolume" => "ask_best_volume"
"askBest30MWPrice" => "ask_best_30MW_price"
...

The first two examples are rather simple and should be covered by a basic snake_case(name_in_camel_case::String) function. The third example would require defining "reserved words" that are ignored in the conversion.

2 Answers

The snakecase function in the python library you referenced looks like this:

def snakecase(string):
    string = re.sub(r"[\-\.\s]", '_', str(string))
    if not string:
        return string
    return lowercase(string[0]) + re.sub(r"[A-Z]", lambda matched: '_' + lowercase(matched.group(0)), string[1:])

snakecase("timeStamp") #time_stamp
snakecase("askBestVolume") #ask_best_volume
snakecase("askBest30MWPrice") #ask_best30_m_w_price

One way to write it in julia is:

function snakecase(string)
    string = replace(string, r"[\-\.\s]" => "_")
    words = lowercase.(split(string, r"(?=[A-Z])"))
    return join([i==1 ? word : "_$word" for (i,word) in enumerate(words)])
end

snakecase("timeStamp") #time_stamp
snakecase("askBestVolume") #ask_best_volume
snakecase("askBest30MWPrice") #ask_best30_m_w_price

Note however that "askBest30MWPrice" is not handled the way you wrote (in either function)

This doesn't use "reserved words" as mentioned in the question, but instead assumes that a series of upper case letters (along with preceding numbers if any, for eg. "30MW") is supposed to be a word of its own; while also ensuring that "Price" in "30MWPrice" is seen as a separate word.


function snake_case(camelstring::S) where S<:AbstractString

  wordpat = r"
  ^[a-z]+ |                  #match initial lower case part
  [A-Z][a-z]+ |              #match Words Like This
  \d*([A-Z](?=[A-Z]|$))+ |   #match ABBREV 30MW 
  \d+                        #match 1234 (numbers without units)
  "x

  smartlower(word) = any(islowercase, word) ? lowercase(word) : word
  words = [smartlower(m.match) for m in eachmatch(wordpat, camelstring)]

  join(words, "_")
end

using Test

function runtests()
  @test snake_case("askBest30MWPrice") == "ask_best_30MW_price"
  @test snake_case("welcomeToAIOverlords") == "welcome_to_AI_overlords"
  @test snake_case("queryInterface") == "query_interface"
  @test snake_case("tst") == "tst"
  @test snake_case("1234") == "1234"
  @test snake_case("column12Value") == "column_12_value"
  @test snake_case("readTOC") == "read_TOC"
end

(Probably Unimportant) Side Note: You can replace [a-z] with [[:lower:]] and [A-Z] with [[:upper:]] above to make it work for some more languages, for eg. snake_case("helloΩorld") will then return "hello_ωorld". However (like with anything Unicode), there are nuances - for eg., my language (Tamil) doesn't have letter cases, so its letters fall outside both [:lower:] and [:upper:].
A proper Unicode solution using \p{Lo}, [:digit:], and whatever else, is left as an exercise to the reader.

Related