You can also see Cython, Java, Python, C++, C, Js, or C# repository.
DataSet yaratmak için AnnotatedDataSetGenerator sınıfı önce üretilir.
AnnotatedDataSetGenerator(folder: string, pattern: string, instanceGenerator: InstanceGenerator)
Ardından generate metodu ile DataSet yaratılır.
generate(): DataSet
DataGeneratorlerin InstanceGeneratorlere ihtiyacı vardır. Bunlar bir tek kelimeden bir Instance yaratan sınıflardır.
generateInstanceFromSentence(sentence: Sentence, wordIndex: number): Instance
NER problemi için NerInstanceGenerator, FeaturedNerInstanceGenerator ve VectorizedNerInstanceGeneratorsınıfı
ShallowParse problemi için ShallowParseInstanceGenerator, FeaturedShallowParseInstanceGenerator ve VectorizedShallowParseInstanceGenerator sınıfı
WSD problemi için SemanticInstanceGenerator, FeaturedSemanticInstanceGenerator ve VectorizedSemanticInstanceGenerator sınıfı
Morphological Disambiguation problemi için FeaturedDisambiguationInstanceGenerator sınıfı
The following Table shows the sample text represented with sense labels and three possible features, namely the root form of the word, the part of speech (POS) tag of the word, and a boolean feature for checking the capital case.
| Word | Root | Pos | Capital | ... | Tag |
|---|---|---|---|---|---|
| Yüzündeki | yüz | Noun | True | ... | yüz3 |
| ketçap | ketçap | Noun | False | ... | ketçap1 |
| lekesi | leke | Noun | False | ... | leke2 |
| yüzdükten | yüz | Verb | False | ... | yüz2 |
| sonra | sonra | PCAbl | False | ... | sonra1 |
| çıkmış | çık | Verb | False | ... | çık10 |
| . | . | Punctuation | False | ... | .1 |
The following Table shows the sample text represented with tag labels and three possible features, namely the root form of the word, the part of speech (POS) tag of the word, and a boolean feature for checking the capital case.
| Word | Root | Pos | Capital | ... | Tag |
|---|---|---|---|---|---|
| Türk | Türk | Noun | True | ... | ORGANIZATION |
| Hava | Hava | Noun | True | ... | ORGANIZATION |
| Yolları | Yol | Noun | True | ... | ORGANIZATION |
| bu | bu | Pronoun | False | ... | NONE |
| Pazartesi'den | Pazartesi | Noun | True | ... | TIME |
| itibaren | itibaren | Adverb | False | ... | NONE |
| İstanbul | İstanbul | Noun | True | ... | LOCATION |
| Ankara | Ankara | Noun | True | ... | LOCATION |
| güzergahı | güzergah | Noun | False | ... | NONE |
| için | için | Adverb | False | ... | NONE |
| indirimli | indirimli | Adjective | False | ... | NONE |
| satışlarını | sat | Noun | False | ... | NONE |
| 90 | 90 | Number | False | ... | MONEY |
| TL'den | TL | Noun | True | ... | MONEY |
| başlatacağını | başlat | Noun | False | ... | NONE |
| açıkladı | açıkla | Verb | False | ... | NONE |
| . | . | Punctuation | False | ... | NONE |
The following Table shows the sample text represented with chunk labels and three possible features, namely the root form of the word, the part of speech (POS) tag of the word, and a boolean feature for checking the capital case.
| Word | Root | Pos | Capital | ... | Tag |
|---|---|---|---|---|---|
| Türk | Türk | Noun | True | ... | ÖZNE |
| Hava | Hava | Noun | True | ... | ÖZNE |
| Yolları | yol | Noun | True | ... | ÖZNE |
| Salı | Salı | Noun | True | ... | ZARF TÜMLECİ |
| günü | gün | Noun | False | ... | ZARF TÜMLECİ |
| yeni | yeni | Adjective | False | ... | NESNE |
| indirimli | indirimli | Adjective | False | ... | NESNE |
| fiyatlarını | fiyat | Noun | False | ... | NESNE |
| açıkladı | açıkla | Verb | False | ... | YÜKLEM |
| . | . | Punctuation | False | ... | HİÇBİRİ |
If you use this resource on your research, please cite the following paper:
@article{acikgoz,
title={All-words word sense disambiguation for {T}urkish},
author={O. Açıkg{\"o}z and A. T. G{\"u}rkan and B. Ertopçu and O. Topsakal and B. {\"O}zenç and A. B. Kanburoğlu and {\.{I}}. Çam and B. Avar and G. Ercan and O. T. Y{\i}ld{\i}z},
journal={2017 International Conference on Computer Science and Engineering (UBMK)},
year={2017},
pages={490-495}
}
@inproceedings{ertopcu17,
author={B. {Ertopçu} and A. B. {Kanburoğlu} and O. {Topsakal} and O. {Açıkgöz} and A. T. {Gürkan} and B. {Özenç} and İ. {Çam} and B. {Avar} and G. {Ercan} and O. T. {Yıldız}},
booktitle={2017 International Conference on Computer Science and Engineering (UBMK)}, title={A new approach for named entity recognition},
year={2017},
pages={474-479}
}
For Contibutors
============
### Package.swift file
1. Dependencies should be given w.r.t github.
dependencies: [
.package(name: "MorphologicalAnalysis", url: "https://github.com/StarlangSoftware/TurkishMorphologicalAnalysis-Swift.git", .exact("1.0.6"))],
2. Targets should include direct dependencies, files to be excluded, and all resources.
targets: [
.target(
dependencies: ["MorphologicalAnalysis"],
exclude: ["turkish1944_dictionary.txt", "turkish1944_wordnet.xml",
"turkish1955_dictionary.txt", "turkish1955_wordnet.xml",
"turkish1959_dictionary.txt", "turkish1959_wordnet.xml",
"turkish1966_dictionary.txt", "turkish1966_wordnet.xml",
"turkish1969_dictionary.txt", "turkish1969_wordnet.xml",
"turkish1974_dictionary.txt", "turkish1974_wordnet.xml",
"turkish1983_dictionary.txt", "turkish1983_wordnet.xml",
"turkish1988_dictionary.txt", "turkish1988_wordnet.xml",
"turkish1998_dictionary.txt", "turkish1998_wordnet.xml"],
resources:
[.process("turkish_wordnet.xml"),.process("english_wordnet_version_31.xml"),.process("english_exception.xml")]),
3. Test targets should include test directory.
.testTarget(
name: "WordNetTests",
dependencies: ["WordNet"]),
### Data files
1. Add data files to the project folder.
### Swift files
1. Do not forget to comment each function.
/** * Returns the value to which the specified key is mapped. - Parameters: - id: String id of a key - Returns: value of the specified key */ public func singleMap(id: String) -> String{ return map[id]! }
2. Do not forget to define classes as open in order to be able to extend them in other packages.
open class Word : Comparable, Equatable, Hashable
3. Function names should follow caml case.
public func map(id: String)->String?
4. Write getter and setter methods.
public func getSynSetId() -> String{
public func setOrigin(origin: String){
5. Use separate test class extending XCTestCase for testing purposes.
final class WordNetTest: XCTestCase { var turkish : WordNet = WordNet()
func testSize() {
XCTAssertEqual(78326, turkish.size())
}
6. Enumerated types should be declared as enum.
public enum CategoryType : String{ case MATHEMATICS case SPORT case MUSIC
7. Implement == operator and hasher method for hashing purposes.
public func hash(into hasher: inout Hasher) {
hasher.combine(name)
}
public static func == (lhs: Relation, rhs: Relation) -> Bool {
return lhs.name == rhs.name
}
8. Make classes Comparable for comparison, Equatable for equality, and Hashable for hashing check.
open class Word : Comparable, Equatable, Hashable
9. Implement < operator for comparison purposes.
public static func < (lhs: Word, rhs: Word) -> Bool {
return lhs.name < rhs.name
}
10. Implement description for toString method.
open func description() -> String{
11. Use Bundle and XMLParserDelegate for parsing Xml files.
let url = Bundle.module.url(forResource: fileName, withExtension: "xml")
var parser : XMLParser = XMLParser(contentsOf: url!)!
parser.delegate = self
parser.parse()
also use parser method.
public func parser(_ parser: XMLParser, didStartElement elementName: String, namespaceURI: String?, qualifiedName qName: String?, attributes attributeDict: [String : String])

