Novotny / Bilokon / Galiotos | Machine Learning and Big Data with Kdb+/Q | Buch | 978-1-119-40475-0 | www.sack.de

Buch, Englisch, 640 Seiten, Format (B × H): 179 mm x 250 mm, Gewicht: 1204 g

Novotny / Bilokon / Galiotos

Machine Learning and Big Data with Kdb+/Q


1. Auflage 2020
ISBN: 978-1-119-40475-0
Verlag: Wiley

Buch, Englisch, 640 Seiten, Format (B × H): 179 mm x 250 mm, Gewicht: 1204 g

ISBN: 978-1-119-40475-0
Verlag: Wiley


Upgrade your programming language to more effectively handle high-frequency data

Machine Learning and Big Data with KDB+/Q offers quants, programmers and algorithmic traders a practical entry into the powerful but non-intuitive kdb+ database and q programming language. Ideally designed to handle the speed and volume of high-frequency financial data at sell- and buy-side institutions, these tools have become the de facto standard; this book provides the foundational knowledge practitioners need to work effectively with this rapidly-evolving approach to analytical trading.

The discussion follows the natural progression of working strategy development to allow hands-on learning in a familiar sphere, illustrating the contrast of efficiency and capability between the q language and other programming approaches. Rather than an all-encompassing “bible”-type reference, this book is designed with a focus on real-world practicality to help you quickly get up to speed and become productive with the language.
- Understand why kdb+/q is the ideal solution for high-frequency data
- Delve into “meat” of q programming to solve practical economic problems
- Perform everyday operations including basic regressions, cointegration, volatility estimation, modelling and more
- Learn advanced techniques from market impact and microstructure analyses to machine learning techniques including neural networks

The kdb+ database and its underlying programming language q offer unprecedented speed and capability. As trading algorithms and financial models grow ever more complex against the markets they seek to predict, they encompass an ever-larger swath of data – more variables, more metrics, more responsiveness and altogether more “moving parts.”

Traditional programming languages are increasingly failing to accommodate the growing speed and volume of data, and lack the necessary flexibility that cutting-edge financial modelling demands. Machine Learning and Big Data with KDB+/Q opens up the technology and flattens the learning curve to help you quickly adopt a more effective set of tools.

Novotny / Bilokon / Galiotos Machine Learning and Big Data with Kdb+/Q jetzt bestellen!

Weitere Infos & Material


Preface xvii

About the Authors xxiii

Part One Language Fundamentals

Chapter 1 Fundamentals of the q Programming Language 3

1.1 The (Not So Very) First Steps in q 3

1.2 Atoms and Lists 5

1.3 Basic Language Constructs 14

1.4 Basic Operators 19

1.5 Difference between Strings and Symbols 31

1.6 Matrices and Basic Linear Algebra in q 33

1.7 Launching the Session: Additional Options 35

1.8 Summary and How-To’s 38

Chapter 2 Dictionaries and Tables: The q Fundamentals 41

2.1 Dictionary 41

2.2 Table 44

2.3 The Truth about Tables 48

2.4 Keyed Tables are Dictionaries 50

2.5 From a Vector Language to an Algebraic Language 51

Chapter 3 Functions 57

3.1 Namespace 59

3.2 The Six Adverbs 60

3.3 Apply 72

3.4 Protected Evaluations 75

3.5 Vector Operations 76

3.6 Convention for User-Defined Functions 79

Chapter 4 Editors and Other Tools 81

4.1 Console 81

4.2 Jupyter Notebook 82

4.3 GUIs 84

4.4 IDEs: IntelliJ IDEA 90

4.5 Conclusion 92

Chapter 5 Debugging q Code 93

5.1 Introduction to Making It Wrong: Errors 93

5.2 Debugging the Code 100

5.3 Debugging Server-Side 102

Part Two Data Operations

Chapter 6 Splayed and Partitioned Tables 107

6.1 Introduction 107

6.2 Saving a Table as a Single Binary File 108

6.3 Splayed Tables 110

6.4 Partitioned Tables 113

6.5 Conclusion 119

Chapter 7 Joins 121

7.1 Comma Operator 121

7.2 Join Functions 125

7.3 Advanced Example: Running TWAP 144

Chapter 8 Parallelisation 151

8.1 Parallel Vector Operations 152

8.2 Parallelisation over Processes 155

8.3 Map-Reduce 155

8.4 Advanced Topic: Parallel File/Directory Access 158

Chapter 9 Data Cleaning and Filtering 161

9.1 Predicate Filtering 161

9.2 Data Cleaning, Normalising and APIs 163

Chapter 10 Parse Trees 165

10.1 Definition 166

10.2 Functional Queries 171

Chapter 11 A Few Use Cases 181

11.1 Rolling VWAP 181

11.2 Weighted Mid for N Levels of an Order Book 183

11.3 Consecutive Runs of a Rule 185

11.4 Real-Time Signals and Alerts 186

Part Three Data Science

Chapter 12 Basic Overview of Statistics 191

12.1 Histogram 191

12.2 First Moments 196

12.3 Hypothesis Testing 198

Chapter 13 Linear Regression 229

13.1 Linear Regression 230

13.2 Ordinary Least Squares 231

13.3 The Geometric Representation of Linear Regression 233

13.4 Implementation of the OLS 240

13.5 Significance of Parameters 243

13.6 How Good is the Fit: R2 244

13.7 Relationship with Maximum Likelihood Estimation and AIC with Small Sample Correction 248

13.8 Estimation Suite 252

13.9 Comparing Two Nested Models: Towards a Stopping Rule 254

13.10 In-/Out-of-Sample Operations 257

13.11 Cross-validation 262

13.12 Conclusion 264

Chapter 14 Time Series Econometrics 265

14.1 Autoregressive and Moving Average Processes 265

14.2 Stationarity and Granger Causality 285

14.3 Vector Autoregression 287

Chapter 15 Fourier Transform 301

15.1 Complex Numbers 301

15.2 Discrete Fourier Transform 308

15.3 Addendum: Quaternions 314

15.4 Addendum: Fractals 321

Chapter 16 Eigensystem and PCA 325

16.1 Theory 325

16.2 Algorithms 327

16.3 Implementation of Eigensystem Calculation 332

16.4 The Data Matrix and the Principal Component Analysis 341

16.5 Implementation of PCA 351

16.6 Appendix: Determinant 354

Chapter 17 Outlier Detection 359

17.1 Local Outlier Factor 360

Chapter 18 Simulating Asset Prices 369

18.1 Stochastic Volatility Process with Price Jumps 369

18.2 Towards the Numerical Example 371

18.3 Conclusion 378

Part Four Machine Learning

Chapter 19 Basic Principles of Machine Learning 381

19.1 Non-Numeric Features and Normalisation 381

19.2 Iteration: Constructing Machine Learning Algorithms 386

Chapter 20 Linear Regression with Regularisation 391

20.1 Bias–Variance Trade-off 392

20.2 Regularisation 393

20.3 Ridge Regression 394

20.4 Implementation of the Ridge Regression 396

20.5 Lasso Regression 403

20.6 Implementation of the Lasso Regression 405

Chapter 21 Nearest Neighbours 419

21.1 k-Nearest Neighbours Classifier 419

21.2 Prototype Clustering 423

21.3 Feature Selection: Local Nearest Neighbours Approach 429

Chapter 22 Neural Networks 437

22.1 Theoretical Introduction 437

22.2 Implementation of Neural Networks 445

22.3 Examples 451

22.4 Possible Suggestions 463

Chapter 23 AdaBoost with Stumps 465

23.1 Boosting 465

23.2 Decision Stumps 466

23.3 AdaBoost 467

23.4 Implementation of AdaBoost 468

23.5 Recommendation for Readers 474

Chapter 24 Trees 477

24.1 Introduction to Trees 477

24.2 Regression Trees 479

24.3 Classification Tree 482

24.4 Miscellaneous 484

24.5 Implementation of Trees 485

Chapter 25 Forests 495

25.1 Bootstrap 495

25.2 Bagging 498

25.3 Implementation 500

Chapter 26 Unsupervised Machine Learning: The Apriori Algorithm 509

26.1 Apriori Algorithm 510

26.2 Implementation of the Apriori Algorithm 511

Chapter 27 Processing Information 523

27.1 Information Retrieval 523

27.2 Information as Features 532

Chapter 28 Towards AI – Monte Carlo Tree Search 541

28.1 Multi-Armed Bandit Problem 541

28.2 Monte Carlo Tree Search 558

28.3 Monte Carlo Tree Search Implementation – Tic-tac-toe 565

28.4 Monte Carlo Tree Search – Additional Comments 579

Chapter 29 Econophysics: The Agent-Based Computational Models 583

29.1 Agent-Based Modelling 584

29.2 Ising Agent-Based Model for Financial Markets 587

29.3 Conclusion 592

Chapter 30 Epilogue: Art 595

Bibliography 601

Index 607


JAN NOVOTNY is an eFX quant trader at Deutsche Bank. Previously, he worked at the Centre for Econometric Analysis on high-frequency econometric models. He holds a PhD from CERGE-EI, Charles University, Prague.

PAUL A. BILOKON is CEO and founder of Thalesians Ltd and an expert in algorithmic trading. He previously worked at Nomura, Lehman Brothers, and Morgan Stanley. Paul was educated at Christ Church College, Oxford, and Imperial College.

ARIS GALIOTOS is the global technical lead for the eFX kdb+ team at HSBC, where he helps develop a big data installation processing billions of real-time records per day. Aris holds an MSc in Financial Mathematics with Distinction from the University of Edinburgh.

FRÉDÉRIC DÉLÈZE is an independent algorithm trader and consultant. He has designed automated trading strategies for hedge funds and developed quantitative risk models for investment banks. He holds a PhD in Finance from Hanken School of Economics, Helsinki.



Ihre Fragen, Wünsche oder Anmerkungen
Vorname*
Nachname*
Ihre E-Mail-Adresse*
Kundennr.
Ihre Nachricht*
Lediglich mit * gekennzeichnete Felder sind Pflichtfelder.
Wenn Sie die im Kontaktformular eingegebenen Daten durch Klick auf den nachfolgenden Button übersenden, erklären Sie sich damit einverstanden, dass wir Ihr Angaben für die Beantwortung Ihrer Anfrage verwenden. Selbstverständlich werden Ihre Daten vertraulich behandelt und nicht an Dritte weitergegeben. Sie können der Verwendung Ihrer Daten jederzeit widersprechen. Das Datenhandling bei Sack Fachmedien erklären wir Ihnen in unserer Datenschutzerklärung.