NLTK word tokenize all words except words with a dash, e.g ('hi-there', 'me-you')

Viewed 1525

I am not sure how can I use the nltk.word_tokenize method if I want to tokenize everything except the words with a dash (i.e excludes all words that have a dash in between). example:

'hi-there', 'me-you'

I have tried using the RegexpTokenizer and writing a regex but I somehow make it fail to act like the word_tokenize method and exclude '-'.

Input: 'hello I am an artificial-human'

Output im looking for:

['hello','I','am','an','artificial-human']
4 Answers

The answer that Jay gives you will separate correctly the words that are connected by a dash but you will have to afterwards use bigram of words in order to learn about these combination of words.

For instance, if you are doing a TF-IDF afterwards you could generate it like this:

TfidfVectorizer(ngram_range = (1,2)) 

This will generate a vectorizer taking into account unigrams and bigrams of words.

You could also replace the dash with empty and just concat the two words in one, to afterwards tokenize the words as one alone and have the dash sepparated word as whole words.

text = text.replace('-', '')
text = nltk.tokenize.word_tokenize(text)

Output:

['hello','I','am','an','artificialhuman']

Here are two ways I suggest.This first one is using split() function.Yes not an ideal choice for tokenising,but is easy and seems to do what you want to get.

print('hello I am an artificial-human'.split())

If you still want to use NLTK, you can use Whitespacetokenizer

t='hello I am an artificial-human'
import nltk
from nltk.tokenize import WhitespaceTokenizer
x=WhitespaceTokenizer().tokenize(t)
print(x)

Output of both cases:

enter image description here

I am not an expert in NLTK as id not know how this tokeniser well behaves in any other situations.I saw this example from this article, take a look if you have some doubts

Here is my solution.

You can partition your main string using a sub-string (for e.g, "artificial-human", Used regex for this) into first, the sub_string and last.

And I tokenize the first and recursively tokenize the last and return everything.

import re
import nltk

def regex_partition(string, regex):

    first = re.split(regex, string, 1)[0]
    try:
        last = re.split(regex, string, 1)[1]
    except IndexError:
        last = ''
    
    regp = re.compile(regex)
    result = regp.search(string)
    
    try:
        match = result.group()
    except AttributeError:
        match = ''
    
    return first, match, last

def my_own_tokenizer(string):
    
    first, sub_string, last = regex_partition(string, "[a-zA-Z]+-[a-zA-Z]+")
    
    if sub_string:
        tokens = my_own_tokenizer(last)
        return nltk.word_tokenize(first) + [sub_string] + tokens
    else:
        return nltk.word_tokenize(first)
In [2]: my_own_tokenizer("hello I am an artificial-human")
Out[2]: ['hello', 'I', 'am', 'an', 'artificial-human']

In [3]: my_own_tokenizer("hello I am an artificial-human, how are you?")
Out[3]: ['hello', 'I', 'am', 'an', 'artificial-human', ',', 'how', 'are', 'you', '?']

In [4]: my_own_tokenizer("artifical-human here how are you?")
Out[4]: ['artifical-human', 'here', 'how', 'are', 'you', '?']

In [5]: my_own_tokenizer("hello-word I am an artifical-human!")
Out[5]: ['hello-word', 'I', 'am', 'an', 'artifical-human', '!']

You could replace all instances of "-" with a space before you process the text:

text = text.replace("-", " ")
text = nltk.tokenize.word_tokenize(text)

Of course, this means any instances where you do want to keep the "-" aren't taken into account (however, I'm not sure if there are any tokenizers that have that behaviour, I can't think of any scenarios where you would want keep the hypen).

If you're willing to switch libraries, spacy is an option that does what you want:

import spacy
nlp = spacy.load("en_core_web_sm")

for token in nlp("hello-world, nice to meet you!"):
    print(token)
hello
-
world
,
nice
to
meet
you
!
Related