explosion / spaCy

💫 Industrial-strength Natural Language Processing (NLP) in Python
https://spacy.io
MIT License
29.81k stars 4.37k forks source link

Training a spacy model for NER in french resumes dont give any results #5279

Closed ghost closed 4 years ago

ghost commented 4 years ago

Sample of trainning data(input.json), the full json has only 100 resumes.

{"content": "Resume 1 text in french","annotation":[{"label":["diplomes"],"points":[{"start":1233,"end":1423,"text":"1995-1996 : Lycée  Dar Essalam                                                                     Rabat     \n                        Baccalauréat scientifique option sciences Expérimentales "}]},{"label":["diplomes"],"points":[{"start":1012,"end":1226,"text":"1996-1998 : Faculté des Sciences                                                                          Rabat \n                  C.E.U.S (Certificat des Etudes universitaires Supérieurs) option physique et chimie "}]},{"label":["diplomes"],"points":[{"start":812,"end":1004,"text":"1999-2000 : Faculté des Sciences                                                                           Rabat \n                            Licence es sciences physique  option électronique  "}]},{"label":["diplomes"],"points":[{"start":589,"end":805,"text":"2002-2004 : Faculté des Sciences                                                                           Rabat  \nDESA ((Diplôme des Etudes Supérieures Approfondies)  en informatique   \n\ntélécommunication multimédia "}]},{"label":["diplomes"],"points":[{"start":365,"end":582,"text":"2014-2017 : Institut National des Postes et Télécommunications INPT                 Rabat                                           \n                             Thèse de doctorat en informatique et télécommunication  "}]},{"label":["adresse"],"points":[{"start":122,"end":157,"text":"Rue 34 n 17 Hay Errachad Rabat Maroc"}]}],"extras":null,"metadata":{"first_done_at":1586140561000,"last_updated_at":1586140561000,"sec_taken":0,"last_updated_by":"wP21IMXff9TFSNLNp5v0fxbycFX2","status":"done","evaluation":"NONE"}}

{"content": "Resume 2 text in french","annotation":[{"label":["diplomes"],"points":[{"start":1251,"end":1345,"text":"Lycée Oued El Makhazine - Meknès \n\n- Bachelier mention très bien \n- Option : Sciences physiques"}]},{"label":["diplomes"],"points":[{"start":1122,"end":1231,"text":"Classes préparatoires Moulay Youssef - Rabat \n\n- Admis au Concours National Commun CNC \n- Option : PCSI - PSI "}]},{"label":["diplomes"],"points":[{"start":907,"end":1101,"text":"Institut National des Postes et Télécommunications INPT - Rabat \n\n- Ingénieur d’État en Télécommunications et technologies de l’information \n- Option : MTE Management des Télécoms de l’entreprise"}]},{"label":["adresse"],"points":[{"start":79,"end":133,"text":"94, Hay El Izdihar, Avenue El Massira, Ouislane, MEKNES"}]}],"extras":null,"metadata":{"first_done_at":1586126476000,"last_updated_at":1586325851000,"sec_taken":0,"last_updated_by":"wP21IMXff9TFSNLNp5v0fxbycFX2","status":"done","evaluation":"NONE"}}

{"content": "Resume 3 text in french","annotation":[{"label":["adresse"],"points":[{"start":2757,"end":2804,"text":"N141 Av. El Hansali Agharass \nBouargane \nAgadir "}]},{"label":["diplomes"],"points":[{"start":262,"end":369,"text":"2009-2010 :  Baccalauréat Scientifique, option : Sciences Physiques au Lycée Qualifiant \nIBN MAJJA à Agadir."}]},{"label":["diplomes"],"points":[{"start":125,"end":259,"text":"2010-2016 :  Diplôme d’Ingénieur d’Etat, option : Génie Informatique, à l’Ecole  \nNationale des Sciences Appliquées d’Agadir (ENSAA).  "}]}],"extras":null,"metadata":{"first_done_at":1586141779000,"last_updated_at":1586141779000,"sec_taken":0,"last_updated_by":"wP21IMXff9TFSNLNp5v0fxbycFX2","status":"done","evaluation":"NONE"}}

{"content": "Resume 4 text in french","annotation":[{"label":["diplomes"],"points":[{"start":505,"end":611,"text":"2012 Baccalauréat Sciences Expérimentales option Sciences Physiques, Lycée Hassan Bno \nTabit, Ouled Abbou. "}]},{"label":["diplomes"],"points":[{"start":375,"end":499,"text":"2012–2015 Diplôme de licence en Informatique et Gestion Industrielle, IGI, Faculté des sciences \net Techniques, Settat, LST. "}]},{"label":["diplomes"],"points":[{"start":272,"end":367,"text":"2015–2017 Master Spécialité BioInformatique et Systèmes Complexes, BISC, ENSA , Tanger, \n\nBac+5."}]},{"label":["adresse"],"points":[{"start":15,"end":71,"text":"246 Hay Pam Eljadid OULED ABBOU  \n26450 BERRECHID, Maroc "}]}],"extras":null,"metadata":{"first_done_at":1586127374000,"last_updated_at":1586327010000,"sec_taken":0,"last_updated_by":"wP21IMXff9TFSNLNp5v0fxbycFX2","status":"done","evaluation":"NONE"}}

Code that transformes this json data to spacy format

input_file="input.json"
training_data = []
lines=[]

with open(input_file, 'r', encoding="utf8") as f:
    lines = f.readlines()

for line in lines:
    data = json.loads(line)
    text = data['content']
    entities = []
    for annotation in data['annotation']:
        point = annotation['points'][0]
        labels = annotation['label']
        if not isinstance(labels, list):
            labels = [labels]

        for label in labels:
            entities.append((point['start'], point['end'] + 1 ,label))

    training_data.append((text, {"entities" : entities}))
print(training_data)

Code for training the spacy model

def train_spacy():
    TRAIN_DATA = training_data
    nlp = spacy.blank('en')
    ner = nlp.create_pipe("ner")
    nlp.add_pipe(ner, last=True)
    ner = nlp.get_pipe("ner")
    # add labels
    for _, annotations in TRAIN_DATA:
         for ent in annotations.get('entities'):
            ner.add_label(ent[2])

    optimizer = nlp.begin_training()
    for itn in range(20):
        losses = {}
        batches = minibatch(TRAIN_DATA, size=compounding(4.0, 32.0, 1.001))
        for batch in batches:
            nlp.update(
                [text],  # batch of texts
                [annotations],  # batch of annotations
                drop=0.1,  # dropout - make it harder to memorise data
                sgd=optimizer,  # callable to update weights
                losses=losses
            )
        print(itn, dt.datetime.now(), losses)

    return nlp

Code for testing the model, I test with a text from the training, so normally it should give my the 2 entities present in it(diplomes, adresse)

nlp = train_spacy()
print(nlp.pipeline)
#Resume text
txt_cv = "Dr.XXXX XXXXXXX                                  \n\n \nEmail  : XXXXX@XXXXXX.com \n\nGSM   : XXXXXXXXXXX \n\nAdresse : Rue 34 n 17 Hay Errachad Rabat Maroc \n \n\n \n\nETAT CIVIL \n \n\nSituation de famille : célibataire  \n\nNationalité              : Marocaine \n\nNé le                        : XX XXX XXXX \n\nLieu de naissance   : Talsinnte figuig \n\n \n FORMATION \n\n• 2014-2017 : Institut National des Postes et Télécommunications INPT                 Rabat                                           \n                             Thèse de doctorat en informatique et télécommunication  \n \n\n• 2002-2004 : Faculté des Sciences                                                                           Rabat  \nDESA ((Diplôme des Etudes Supérieures Approfondies)  en informatique   \n\ntélécommunication multimédia \n \n\n• 1999-2000 : Faculté des Sciences                                                                           Rabat \n                            Licence es sciences physique  option électronique  \n \n\n•  1996-1998 : Faculté des Sciences                                                                          Rabat \n                  C.E.U.S (Certificat des Etudes universitaires Supérieurs) option physique et chimie \n \n\n• 1995-1996 : Lycée  Dar Essalam                                                                     Rabat     \n                        Baccalauréat scientifique option sciences Expérimentales \n\nSTAGE  DE FORMATION \n\n• Du 03/03/2004  au 17/09/2004 : Stage de Projet de Fin d’Etudes à l’ INPT  pour  \nl’obtention du  DESA                (Diplôme des Etudes Supérieures Approfondies). \n\n                                  Sujet : AGENT RMON DANS LA GESTION DE RESEAUX. \n\n• Du 03/06/2002  au 17/01/2003: Stage de Projet de Fin d’année à INPT \n  Sujet : Mécanisme d’Authentification Kerbéros Dans un Réseau Sans fils sous Redhat. \n\nPUBLICATION  \n\n✓ Ababou, Mohamed, Rachid Elkouch, and Mostafa Bellafkih and Nabil Ababou. \"New \n\nstrategy to optimize the performance of epidemic routing protocol.\" International Journal \n\nof Computer Applications, vol. 92, N.7, 2014.  \n\n✓ Ababou, Mohamed, Rachid Elkouch, and Mostafa Bellafkih and Nabil Ababou. \"New \n\nStrategy to optimize the Performance of Spray and wait Routing Protocol.\" International \n\nJournal of Wireless and Mobile Networks v.6, N.2, 2014. \n\n✓ Ababou, Mohamed, Rachid Elkouch, and Mostafa Bellafkih and Nabil Ababou. \"Impact of \n\nmobility models on Supp-Tran optimized DTN Spray and Wait routing.\" International \n\njournal of Mobile Network Communications & Telematics ( IJMNCT), Vol.4, N.2, April \n\n2014. \n\n✓ M. Ababou, R. Elkouch, M. Bellafkih and N. Ababou, \"AntProPHET: A new routing \n\nprotocol for delay tolerant networks,\" Proceedings of 2014 Mediterranean Microwave \n\nSymposium (MMS2014), Marrakech, 2014, IEEE. \n\nmailto:mohamed.ababou@gmail.com\n\n\n✓ Ababou, Mohamed, et al. \"BeeAntDTN: A nature inspired routing protocol for delay \n\ntolerant networks.\" Proceedings of 2014 Mediterranean Microwave Symposium \n\n(MMS2014). IEEE, 2014. \n\n✓ Ababou, Mohamed, et al. \"ACDTN: A new routing protocol for delay tolerant networks \n\nbased on ant colony.\" Information Technology: Towards New Smart World (NSITNSW), \n\n2015 5th National Symposium on. IEEE, 2015. \n\n✓ Ababou, Mohamed, et al. \"Energy-efficient routing in Delay-Tolerant Networks.\" RFID \n\nAnd Adaptive Wireless Sensor Networks (RAWSN), 2015 Third International Workshop \n\non. IEEE, 2015. \n\n✓ Ababou, Mohamed, et al. \"Energy efficient and effect of mobility on ACDTN routing \n\nprotocol based on ant colony.\" Electrical and Information Technologies (ICEIT), 2015 \n\nInternational Conference on. IEEE, 2015. \n\n✓ Mohamed, Ababou et al. \"Fuzzy ant colony based routing protocol for delay tolerant \n\nnetwork.\" 10th International Conference on Intelligent Systems: Theories and Applications \n\n(SITA). IEEE, 2015. \n\nARTICLES EN COURS DE PUBLICATION \n\n✓ Ababou, Mohamed, Rachid Elkouch, and Mostafa Bellafkih and Nabil Ababou.”Dynamic \n\nUtility-Based Buffer Management Strategy for Delay-tolerant Networks. “International \n\nJournal of Ad Hoc and Ubiquitous Computing, 2017. ‘accepté par la revue’ \n\n✓ Ababou, Mohamed, Rachid Elkouch, and Mostafa Bellafkih and Nabil Ababou. \"Energy \n\nefficient routing protocol for delay tolerant network based on fuzzy logic and ant colony.\" \n\nInternational Journal of Intelligent Systems and Applications (IJISA), 2017. ‘accepté par la \n\nrevue’ \n\nCONNAISSANCES EN INFORMATIQUE \n\n  \n\nLANGUES \n\nArabe,  Français, anglais. \n\nLOISIRS ET INTERETS PERSONNELS \n\n \n\nVoyages, Photographie, Sport (tennis de table, footing), bénévolat. \n\nSystèmes :  UNIX, DOS, Windows  \n\nLangages :  Séquentiels ( C, Assembleur), Requêtes (SQL), WEB (HTML, PHP, MySQL, \n\nJavaScript), Objets (C++, DOTNET,JAVA) , I.A. (Lisp, Prolog) \n\nLogiciels :  Open ERP (Enterprise Resource Planning), AutoCAD, MATLAB, Visual \n\nBasic, Dreamweaver MX. \n\nDivers :  Bases de données, ONE (Opportunistic Network Environment), NS3,  \n\nArchitecture réseaux,Merise,... \n\n"
doc = nlp(txt_cv)
print(doc.ents)

The test give me an empty array, my goal is to extract addresses(adresse) and academic degrees(diplomes) from a resume.

svlandeg commented 4 years ago

Hi! We try to keep this issue tracker focused specifically on bug reports and feature requests, so I suggest to follow up on the Stackoverflow issue you posted instead.

lock[bot] commented 4 years ago

This thread has been automatically locked since there has not been any recent activity after it was closed. Please open a new issue for related bugs.

lock[bot] commented 4 years ago

This thread has been automatically locked since there has not been any recent activity after it was closed. Please open a new issue for related bugs.