I have a list of texts like this:
inp = """Something at the beginning
References
1. Ryff, C.D. (2014) Psychological Well-Being Revisited: Advances in the Science and Practice of Eudaimonia.
2. Deci, E.L. & Ryan, R.M. (2002) Self-determination research: reflections and future directions.
3. Acedo, F. J., & Casillas, J. C. (2005). Current paradigms in the international management field.
Other References
1. Tarelli, E. (2003), “How to transfer responsibilities from expatriates to local nationals”.
2. Riusala, K. and Suutari, V. (2004), “International knowledge transfers through expatriates”.
3. Wallace, J. (2001), “The benefits of mentoring for female lawyers”.
Something at the end
12. Wallace, J. (2001), “The benefits of mentoring for female lawyers”.
Something else at the end"""
The 'Other References' part is present in some texts, in others, the next part begins with, let's say, 'Good References'. Also, similar strings could appear anywhere in the texts. All reference strings are sometimes separated by '\n', sometimes separated just by spaces. Also, '\n' could occur anywhere in the text, just in the middle of reference strings.
I need regex to use in re.findall and return all strings after 'References' in a list of strings like this:
['Ryff, C.D. (2014) Psychological Well-Being Revisited: Advances in the Science and Practice of Eudaimonia.', 'Deci, E.L. & Ryan, R.M. (2002) Self-determination research: reflections and future directions.', 'Acedo, F. J., & Casillas, J. C. (2005). Current paradigms in the international management field.']
But ONLY after 'References' and NOT anywhere earlier or later in the text.
I have been suggested to use this regex:
refs = re.findall(r'^References\s+((?:\d+\.\s*.*?\n)+)', inp, flags=re.M|re.S)
data = ''.join(refs)
output = re.findall(r'\d+\.\s*(.*?)\n', data)
print(output)
But it works only when reference strings are separated by '\n' which is not the case in some texts. And also '\n' could occur anywhere in the texts. I do not need these '\n' at all, so they could be removed from texts.
Example when suggested regex is not working:
inp = """Something at the beginning
References 1. Ryff, C.D. (2014) Psychological Well-Being Revisited: Advances in the Science and Practice of Eudaimonia. Additional Fields. 2. Deci, E.L. & Ryan, R.M. (2002) Self-determination research: reflections and future directions. 3. Acedo, F. J., & Casillas, J. C. (2005). Current paradigms in the international management field. Other References 1. Tarelli, E. (2003), “How to transfer responsibilities from expatriates to local nationals”.
2. Riusala, K. and Suutari, V. (2004), “International knowledge transfers through expatriates”.
3. Wallace, J. (2001), “The benefits of mentoring for female lawyers”.
Something at the end
12. Wallace, J. (2001), “The benefits of mentoring for female lawyers”.
Something else at the end"""
Could anybody please suggest a code that helps me get the list of references?