I have a database containing normative acts (laws, decrees, ordinances, processes, etc). This database is in .txt format and does not follow a structuring pattern.
PORTARIA Nº 392, DE 9 DE SETEMBRO DE 2021
The texts are completely unformatted, with many spaces and line breaks that make their analysis difficult.
The vast majority of these acts, in your body, make reference to other acts such as:
O SECRETÁRIO DE DEFESA AGROPECUÁRIA DO MINISTÉRIO DA AGRICULTURA, PECUÁRIA E ABASTECIMENTO, no uso da atribuição que lhe conferem os art. 21 e 63 do Anexo I do Decreto nº 10.253, de 20 de fevereiro de 2020, tendo em vista o disposto na Lei nº 8.171, de 17 de janeiro de 1991, na Lei nº 13.874, de 20 de setembro de 2019, no Decreto nº 10.178, de 18 de dezembro de 2019, e o que consta do Processo SEI nº 21000.058030 /2020-37, resolve:...
I need to extract from each act the title and its references (if any) and create a CSV file containing two columns. The first always brings the act under analysis and the second its references. Example using the snippet above:
| PORTARIA, 392 | Decreto, 10.253 |
| PORTARIA, 392 | Lei, 8.171 |
| PORTARIA, 392 | Lei, 13.874 |
| PORTARIA, 392 | Decreto, 10.178 |
| PORTARIA, 392 | Processo SEI, 21000.058030 /2020-37 |-----> NOT!
| PORTARIA, 392 | Revogado Lei, 14.133 |
The pattern I observed for the possible accomplishment of the task is the character sequence "Nº".
I checked with a legal professional about the comments made. The SEI Process will not enter the file that will be generated, as it makes no sense. As for the other acts that appear throughout the text MOST follows the pattern act name n° 0,000 or 00,000. He also told me that it was important to bring information such as Revoked by Law No. 14.133 The act that comes before Revoked
Does anyone have an idea of how I can be performing this type of task?
The way I found to perform the task so far was using the script below. But I still don't know how to export structured to CSV.
>>> txt = "O SECRETÁRIO DE DEFESA AGROPECUÁRIA DO MINISTÉRIO DA AGRICULTURA, PECUÁRIA E ABASTECIMENTO, no uso da atribuição que lhe conferem os art. 21 e 63 do Anexo I do Decreto nº 10.253, de 20 de fevereiro de 2020, tendo em vista o disposto na Lei nº 8.171, de 17 de janeiro de 1991, na Lei nº 13.874, de 20 de setembro de 2019, no Decreto nº 10.178, de 18 de dezembro de 2019, e o que consta do Processo SEI nº 21000.058030 /2020-37, resolve:.O SECRETÁRIO DE DEFESA AGROPECUÁRIA DO MINISTÉRIO DA AGRICULTURA, PECUÁRIA E ABASTECIMENTO, no uso da atribuição que lhe conferem os art. 21 e 63 do Anexo I do Decreto nº 10.253, de 20 de fevereiro de 2020, tendo em vista o disposto na Lei nº 8.171, de 17 de janeiro de 1991, na Lei nº 13.874, de 20 de setembro de 2019, no Decreto nº 10.178, de 18 de dezembro de 2019, e o que consta do Processo SEI nº 21000.058030 /2020-37, resolve:."
>>> import re
>>> r = re.findall(r'(Decreto|Lei|Processo SEI) nº (\d+(?: ?[./-]\d+)*)', txt)
>>> print(r)
[('Decreto', '10.253'), ('Lei', '8.171'), ('Lei', '13.874'), ('Decreto', '10.178'), ('Processo SEI', '21000.058030 /2020-37'), ('Decreto', '10.253'), ('Lei', '8.171'), ('Lei', '13.874'), ('Decreto', '10.178'), ('Processo SEI', '21000.058030 /2020-37')]
Thanks!!!